SYSTEM Cited by 12 sources
AWS EKS (Elastic Kubernetes Service)¶
Amazon EKS (Elastic Kubernetes Service) is AWS's managed Kubernetes control plane — AWS runs the API server, etcd, and the core controllers; customers run worker nodes (on EC2 or Fargate) and their own workloads.
From a system-design posture, EKS is the managed-control-plane equivalent of self-hosted Kubernetes with the same data-plane abstractions (Pods, Services, StatefulSets, Ingress, etc.), the same xDS / API / Helm ecosystem, and the same CNCF toolbox (systems/karpenter, systems/keda, systems/envoy, systems/kyverno, systems/cilium).
Stub page — minimal viable for the Figma ECS→EKS migration ingest. Expand on future EKS-internals sources.
Contrast with ECS¶
Figma's 2024 migration post enumerates the EKS advantages that drove their ECS→EKS cutover:
- StatefulSets for stateful workloads — Kubernetes primitive that gives pods stable network identities across restarts. ECS doesn't have this; Figma had written custom container-startup code to dynamically update etcd cluster membership, which was "fragile and hard to maintain." StatefulSets is the standard way to run etcd on Kubernetes.
- Helm charts ecosystem — easy install / upgrade of OSS software (Figma specifically called out systems/temporal). On ECS, the equivalent required hand-porting each service to systems/terraform.
- Graceful node cordon-and-drain. Cordoning a bad EC2 node on EKS lets the API server move pods off respecting shutdown hooks. ECS on EC2 has no equivalent.
- CNCF auto-scaling — systems/keda for pod-level (with custom metrics like SQS queue depth), systems/karpenter for node-level. ECS has some auto-scaling but the CNCF offerings are more flexible.
- Service-mesh availability — Istio (Envoy-based) is trivial to adopt on EKS; on ECS, building equivalent functionality (custom filters, mTLS) would require building in-house what Istio ships.
- Vendor-agnostic user base drives more external investment than ECS (AWS-only).
Seen in¶
- sources/2026-09-30-aws-running-multi-day-az-evacuation-drills-with-arc-zonal-shift
— EKS has native ARC Zonal
Shift support that handles both infra and Kubernetes networking. Enable once
(
eks update-cluster-config --zonal-shift-config enabled=true); on a shift ARC cordons all nodes in the impacted AZ (no new scheduling), removes that AZ's pod endpoints from EndpointSlice resources (east-west traffic redirects to healthy AZs), suspends AZ rebalancing + updates ASGs to launch only in healthy AZs, and preserves nodes/pods in the shifted AZ (not terminated) so full capacity is instant on restore. A full drill shifts both the NLB (north-south) and the cluster ARN (east-west). Caveats: the shift does not evict pods, trigger autoscaling, or control pod→external DB connections — verifytopologySpreadConstraints maxSkew:1ontopology.kubernetes.io/zone, pre-scale for N-1 (static stability), and use Istio locality-aware routing for AZ-specific reader endpoints (Aurora's cluster endpoint already auto-routes to the writer). Restore: cancel the EKS shift before the LB shift. (Source: sources/2026-09-30-aws-running-multi-day-az-evacuation-drills-with-arc-zonal-shift) - sources/2026-08-19-aws-how-clario-detects-phi-pii-in-dicom-images-using-bedrock — EKS runs Clario's PHI/PII detection backend. Rationale: a single DICOM series can span thousands of slices, so the detection workload is long-running and memory-intensive — a poor fit for short-lived function runtimes. AWS SA guidance covered pod-scaling strategy to handle thousands of slices per request without timeout or memory pressure. EKS sits behind API Gateway (which owns auth/rate-limiting) so the backend stays focused on detection.
- sources/2024-08-08-figma-migrated-onto-k8s-in-less-than-12-months — Figma's target platform. Three active EKS clusters per environment receive real traffic for every service — multi-cluster-active-active-redundancy reduces the blast radius of cluster-scoped incidents (like the CoreDNS destruction they describe) to ~1/3 of traffic.
- sources/2026-02-05-aws-convera-verified-permissions-fine-grained-authorization
— EKS as the backend compute tier in Convera's multi-tenant
SaaS flow: API Gateway forwards requests to Kubernetes pods with
tenant_idin a custom header; each pod re-validates with AVP against the tenant's policy store before building a tenant context and forwarding to RDS. EKS pods are the site of the second authorization check in Convera's zero-trust chain. - sources/2026-02-26-aws-santander-catalyst-platform-engineering — EKS as the internal-developer-platform control plane cluster in Santander Catalyst — "the brain of the operation, orchestrating all components and workflows". One EKS cluster hosts three load-bearing sub-components: ArgoCD for data-plane claims (GitOps), OPA Gatekeeper for the policies catalog (policy-gate-on-provisioning), and Crossplane for the stacks catalog (crossplane-composition). This is EKS used as an infrastructure control plane, not an application compute tier — a fundamentally different role from Figma (app compute) or Convera (backend zero-trust tier), and the canonical wiki instance of EKS as the substrate for a multi-cloud internal developer platform.
- sources/2026-03-18-aws-ai-powered-event-response-for-amazon-eks — EKS as the investigation target of AWS DevOps Agent. One Agent Space per EKS cluster; the agent combines a Kubernetes API resource scan (the graph nodes: Pods / Deployments / Services / ConfigMaps / Ingress / NetworkPolicies with their metadata, resource specs, and health checks) with OpenTelemetry-derived runtime relationships (the graph edges: service-mesh traffic, distributed traces, metric attribution) into a unified dependency graph used for root-cause analysis. See telemetry-based-resource-discovery for the discovery methodology and systems/aws-devops-agent for the full investigation workflow.
- sources/2026-03-23-aws-generali-malaysia-eks-auto-mode — EKS in its Auto Mode variant at Generali Malaysia: AWS operates the K8s data plane as well (Bottlerocket AMI on a weekly-replacement cadence, default add-ons, cluster-version upgrades). Canonical wiki reference for the peer-AWS-service integration surface of EKS — the case study documents six managed services plugged into one cluster: GuardDuty (threat detection), Inspector (vuln scanning with ECR-to- running-container mapping), Network Firewall (SNI egress allow-list), Secrets Manager
- External Secrets Operator (env-var secret injection, no volume mounts), Amazon Managed Grafana (per-namespace dashboards), and AWS Billing's split cost allocation data for EKS (eks-cost-allocation-tags). Compound operating discipline: stateless-only pods + immutable pods + Helm-as-standard-packaging + HPA auto-scaling. Customer-retained safety contract under Auto Mode's platform-driven node churn: Pod Disruption Budgets + Node Disruption Budgets + off-peak maintenance window. See systems/generali-malaysia-eks for the full platform synthesis.
- sources/2026-04-06-aws-unlock-efficient-model-deployment-simplified-inference-operator-setup-on-amazon-sagemaker-hyperpod
— EKS as the Kubernetes control plane under SageMaker
HyperPod inference, and (more generally) as the packaging
substrate for the EKS add-on primitive. 2026-04-06
repackaging of the
HyperPod Inference Operator from Helm chart to native EKS
add-on is the canonical wiki instance of
eks-add-on-as-lifecycle-packaging — four
dependency add-ons bundled (cert-manager, S3 Mountpoint CSI,
FSx CSI, metrics-server), four IAM roles scaffolded (execution,
JumpStart gated, ALB controller, KEDA), migration script
(
helm_to_addon.sh) with auto-discovery + OVERWRITE install + rollback semantics. Highlights the EKS add-on API (aws eks create-addon --configuration-values) as a managed-lifecycle packaging primitive that sits alongside Helm as a distribution path for AWS-authored K8s operators. - sources/2026-01-12-aws-salesforce-karpenter-migration-1000-eks-clusters
— EKS as the 1,000-cluster / 1,180-node-pool production
platform at Salesforce — the largest
documented EKS fleet in the wiki. Canonical wiki reference for
EKS-at-extreme-scale operations: Karpenter
migration off Cluster Autoscaler +
ASGs with in-house
transition tool (zero-disruption + PDB-respecting drain +
rollback-to-ASG + CI/CD-integrated); automated ASG→
NodePool/EC2NodeClassconfig mapping over 1,180+ node pools; the five generalisable operational lessons ( PDB hygiene with OPA-enforced admission, sequential node cordoning with verification checkpoints, 63-character label limit as migration-blocker, ** singleton-workload protection under bin-packing consolidation, 1:1 ephemeral-storage translation**). Outcome metrics: scaling latency minutes → seconds; 80% manual-ops reduction; 5% FY2026 cost savings (+5-10% projected for FY2027); eliminated thousands of node groups; heterogeneous GPU / ARM / x86 in single node pools. Rollout: mid-2025 → early 2026, phased with soak times under risk-based sequencing. -
sources/2025-12-11-aws-architecting-conversational-observability-for-cloud-applications — EKS as the investigation target in a self-built AI troubleshooting blueprint, companion to the later AWS-managed DevOps Agent post. Same target, different vendor relationship: a customer-built RAG chatbot (Fluent Bit → Kinesis → Lambda + Bedrock embeddings → OpenSearch Serverless) or Strands-based agent system (with EKS MCP Server for cluster operations) investigates an EKS cluster via an in-cluster troubleshooting assistant pod running with a read-only RBAC service account and a static kubectl allowlist (patterns/tool-surface-minimization). Combined stored telemetry (patterns/telemetry-to-lakehouse) + live
kubectloutput drives an iterative LLM ↔ cluster loop until the LLM judges enough context for resolution. Framing asserts ECS and Lambda as equally valid fabrics for the same approach, though only EKS is demonstrated. -
sources/2026-04-27-aws-deloitte-optimizes-eks-environment-provisioning-with-vcluster — EKS as the shared-host-cluster substrate for 50+ virtual Kubernetes clusters, a role not previously canonicalised on this page. Deloitte runs one EKS cluster with EKS Auto Mode as the host for vCluster-partitioned QA testing environments, each virtual cluster acting like an independent K8s environment while sharing the host's compute + controllers + monitoring. Environment provisioning dropped from 45 min to <5 min (89% reduction), >50 vCPU + >200 GB RAM saved at peak from non-duplicated shared controllers, up to 70% additional savings from EC2 Spot via Auto Mode. See shared-host-cluster-with-virtual-clusters for the topology and shared-alb-path-based-multi-cluster-routing for the companion ingress design that collapses 50+ ALBs into one. This is the first wiki canonical instance of EKS as a vcluster host — contrast with Generali's "EKS Auto Mode as stateless app compute" and Salesforce's "1,000+ EKS clusters under Karpenter" to see the three distinct EKS-as-substrate shapes ingested to date.
-
sources/2026-08-13-aws-adobe-firefly-simplified-observability-with-amazon-managed-prometheus — EKS as the GPU-training telemetry substrate for Adobe Firefly. The metric model spans job-level GPU utilisation, memory, and network throughput; pod and node health; resource allocation; and GPU-health signals. In the source's representative workload, 2,000 nodes and 16,000 GPUs are scraped every 30 seconds, producing more than one billion data points in a query window. AMP managed scrapers collect a critical metric tier alongside the retained self-managed Prometheus stack. This is EKS as the substrate for high-cardinality GPU observability, distinct from its roles as application compute, an infrastructure control plane, or an AI-troubleshooting target. (Source: sources/2026-08-13-aws-adobe-firefly-simplified-observability-with-amazon-managed-prometheus)
EKS's role axis across ingested sources¶
Same platform, substantially different roles per case study:
| Customer | EKS's role |
|---|---|
| Figma | Application compute tier (multi-cluster active-active) |
| Convera | Backend zero-trust compute tier (per-pod AVP reval) |
| Santander | Infrastructure control plane (ArgoCD + OPA + Crossplane) |
| Generali | Multi-tenant app compute tier under Auto Mode |
| SageMaker HyperPod | LLM-inference-platform substrate (EKS add-on packaging) |
| Conversational observability blueprint | AI-troubleshooting target (self-built RAG + MCP variants) |
| AWS DevOps Agent | AI-troubleshooting target (AWS-managed variant) |
| Salesforce | Extreme-scale multi-tenant platform (1,000+ clusters, 1,180+ node pools, Karpenter-driven) |
| Deloitte | Shared host cluster for vCluster virtual clusters (50+ QA environments on 1 EKS cluster) |
| Equinix | Centrally-governed shared-services platform (Workloads/Platform/Network multi-account split; migrated off self-managed K8s) |
This spread is what makes EKS a load-bearing canonical node in the wiki — the same primitive reappears in very different architectures.
Related¶
- systems/kubernetes — what EKS runs
- systems/eks-auto-mode — the managed-data-plane variant
- systems/bottlerocket — the AMI under Auto Mode
- systems/amazon-ecs — the AWS orchestrator EKS is compared with
- systems/karpenter, systems/keda — the CNCF auto-scaling projects that motivated Figma's migration
- systems/crossplane, systems/argocd, systems/open-policy-agent — the CNCF trio Catalyst runs on top of EKS to turn it into an internal-developer-platform control plane
- systems/amazon-guardduty, systems/amazon-inspector, systems/aws-network-firewall, systems/external-secrets-operator, systems/amazon-managed-grafana — the peer AWS services documented as integration surface at Generali
Seen in — hybrid cloud orchestration (2026-09-01)¶
The EKS Distro (EKS-D) that powers managed EKS in the cloud also powers Amazon EKS Anywhere, which runs the entire cluster on customer on-premises hardware (bare-metal provider). The AWS hybrid-cloud orchestration reference architecture manages EKS Anywhere clusters at scale from AWS; EKS Hybrid Nodes (managed cloud control plane + on-prem worker nodes) is the recommended alternative when connectivity to a Region is reliable. (Source: sources/2026-09-01-aws-hybrid-cloud-orchestration-modernizing-on-premises-infrastructure)
Seen in — Equinix shared-services migration (2026-09-17)¶
Equinix migrated off a self-managed Kubernetes estate (EC2-hosted etcd, control-plane, and worker nodes provisioned per app team) onto EKS to escape "operational sprawl" — independently-owned clusters with no shared infra layer, no consistent governance, per-team reinvention of CI/CD / observability / networking, and compounding control-plane upgrade/patch risk. Here EKS is the substrate for a centrally-governed shared-services platform split across three AWS accounts:
- Workloads account — app-team services on EKS across two AZs (us-west-1), ingress via ALB + Gateway API, Cilium CNI for pod-level isolation, Hubble for flow observability.
- Platform account — cloud-ops-owned shared data services (RDS, MSK, OpenSearch, Amazon MQ, S3) consumed cross-account, plus GitHub Runners as EKS workloads for centralized CI/CD; ingress via NLB.
- Network account — Transit Gateway hub across us-west-1 + us-east-2 (inter-Region peering), Route 53 Resolver + private hosted zone for discovery, Direct Connect (dual circuits) to on-prem through network firewalls.
The direct payoff of the managed control plane — AWS owning upgrades, patching, and HA — is a claimed 40% operational-overhead reduction, alongside 4x deployment frequency and weeks → days onboarding via pre-configured self-service namespaces. This extends the "shared platform" role already documented for Santander (control plane) and Deloitte (vCluster host) with a multi-account governed-platform shape. (Source: sources/2026-09-17-aws-how-equinix-cut-operational-overhead-with-a-shared-services)
Multi-tenant isolation reference: ReadyOn's "Four Walls"¶
ReadyOn's Harmony platform (Fortune 100 workforce intelligence) runs multi-tenant on EKS and documents a four-layer tenant-isolation model — a strong worked example of defense in depth on EKS. The premise: a Kubernetes namespace is not a security boundary, so isolation is layered at four independent abstraction layers so that crossing a tenant boundary requires simultaneously defeating all of them:
- Namespace — per-tenant RBAC + admission control, generated identically for every tenant from Git via an Argo CD ApplicationSet (no legacy tenants with weaker policies); GitOps self-heals drift within seconds.
- Compute — Karpenter dedicated per-tenant node pools using a dual-taint strategy (tenant taint + workload-type taint; a pod must tolerate both), with an admission controller validating that tolerations match the namespace's tenant identity. IMDSv2 + reduced hop limit + per-tenant node IAM role cap container-escape blast radius.
- Network — per-tenant VPC security groups (Aurora reachable only from that tenant's node SG), default-deny inter-namespace network policies, and a private API endpoint (VPN-only with OIDC + MFA). The VPC is split into four tiers (public / application / database / control-plane).
- Data — dedicated per-tenant Aurora clusters (no row-level filter to forget), per-tenant KMS keys, and IRSA short-lived STS credentials via OIDC federation between the EKS cluster and IAM.
Every container runs Pod Security Standards restricted (runAsNonRoot,
readOnlyRootFilesystem, capabilities.drop: [ALL], seccompProfile:
RuntimeDefault). ReadyOn maps defenses onto a 10-stage MITRE ATT&CK threat
model and validates that all six cross-tenant paths at the lateral-movement
stage are blocked. This is the canonical wiki example of EKS multi-tenancy
where namespace + scheduler + SDN + IAM are composed as overlapping
controls rather than relying on the namespace alone.
(Source: sources/2026-09-18-aws-readyons-four-walls-of-tenant-isolation-on-amazon-eks)