How Equinix cut operational overhead with a shared services architecture on Amazon EKS¶
Summary¶
Equinix — the digital-infrastructure company running 260+ data centers across 70+ metros — migrated off a self-managed Kubernetes estate (EC2-hosted etcd/control-plane/worker nodes, provisioned per app team) onto a centrally-governed shared services architecture on Amazon EKS. The old model let each application team stand up and own its own cluster, producing "operational sprawl": duplicated infrastructure, no consistent governance mechanism, per-team reinvention of CI/CD, observability, networking, and data services, and compounding control-plane upgrade/patch risk across many self-managed clusters. The new "North Star" architecture is a three-account, multi-AZ EKS design — a Workloads account (app services), a Platform account (cloud-ops-owned shared data + CI/CD services), and a Network account (an AWS Transit Gateway hub spanning two Regions with Direct Connect hybrid links back to on-prem). The cloud operations team centrally owns and governs infrastructure while app teams consume pre-configured namespaces via self-service. Reported results: 40% reduction in operational overhead, 4x increase in deployment frequency, and 100% adoption across multiple business organizations.
Key takeaways¶
- The failure mode was decentralized cluster ownership, not Kubernetes itself. Each app team independently provisioning and running its own self-managed cluster produced sprawl with "no shared infrastructure layer" and "no consistent mechanism to enforce governance, standardize configurations, or provide common infrastructure services." The fix was an organizational + architectural shift to a governed shared platform — a textbook self-service infrastructure / platform-engineering move. (Source: this article)
- Managed control plane removes undifferentiated ops. Moving from self-managed etcd/control-plane/worker EC2 nodes to EKS means "AWS now handles upgrades, patching, and high availability automatically" — the direct source of the claimed 40% operational-overhead reduction. This is the managed-control-plane posture EKS is built for: the compounding upgrade/patch risk of many self-managed control planes disappears.
- Three-account topology encodes isolation and ownership boundaries. The design cleanly separates concerns via account boundaries: Workloads (app-team services), Platform (cloud-ops-owned shared services), and Network (connectivity hub). Multi-account isolation between application and shared-services workloads "reduced the scope of potential security incidents" — tenant isolation and blast-radius reduction at the account layer.
- Cilium CNI enforces pod-level network policy; Hubble gives network observability. Cilium is the Container Network Interface in both the Workloads and Platform accounts, "enforcing network policies that isolate workloads at the pod level," while Hubble provides "real-time network observability across all application traffic flows" — unifying what had been fragmented, team-specific monitoring (observability).
- Ingress differs by account role: ALB + Gateway API for apps, NLB for platform. The Workloads account fronts traffic with an Application Load Balancer plus the Kubernetes Gateway API (GatewayClass/Gateway resources) for "precise routing to application namespaces"; the Platform account uses a Network Load Balancer for its shared services, with the same Cilium/Hubble networking + observability stack.
- Managed data services are consumed across the account boundary. The Platform account houses Amazon RDS, Amazon MSK, Amazon OpenSearch Service, Amazon MQ, and Amazon S3, "consumed by application workloads across the account boundary" — a shared data tier rather than per-team reinvention.
- CI/CD is centralized as GitHub Runners running as EKS workloads. The Platform account runs GitHub Runners as EKS workloads, "providing a centralized, self-service pipeline for all application teams" — standardized CI/CD replacing per-team pipelines, letting app teams deploy independently without cloud-ops intervention.
- The Network account is a multi-Region Transit Gateway hub with hybrid Direct Connect. A dedicated Network account runs an AWS Transit Gateway architecture across us-west-1 and us-east-2 with TGW peering between them; TGW attachments connect the Workloads and Platform VPCs to the central hub; Route 53 Resolver endpoints (inbound + outbound per AZ) plus a private hosted zone provide service discovery; and Direct Connect Gateway with dual circuits connects back to on-prem Equinix border routers through network firewalls for security enforcement.
- Onboarding time collapsed from weeks to days. New application teams onboard "in days rather than weeks," deploying into pre-configured namespaces with access to shared data services, CI/CD, and observability — "without provisioning or managing cluster infrastructure." The self-service platform is the mechanism behind the 4x deployment-frequency gain.
Systems seen¶
- Amazon EKS — managed Kubernetes control plane; the foundation of the North Star architecture. Replaced self-managed etcd/control-plane/worker EC2 nodes.
- AWS Transit Gateway — regional hub-and-spoke router; here a two-Region (us-west-1 / us-east-2) hub with inter-Region peering, connecting Workloads + Platform VPCs via attachments.
- AWS Direct Connect — Direct Connect Gateway with dual circuits to on-prem Equinix border routers (hybrid connectivity).
- Amazon Route 53 — Resolver inbound/outbound endpoints per AZ + private hosted zone for cross-platform service discovery.
- Network Load Balancer — ingress for the Platform account's shared services.
- Application Load Balancer — ingress for the Workloads account, paired with the Gateway API.
- Kubernetes Gateway API — GatewayClass + Gateway resources for precise routing to application namespaces.
- Cilium — CNI enforcing pod-level network isolation in both accounts.
- Hubble — Cilium's network-flow observability layer (real-time visibility across traffic flows); recorded as prose/tag, not minted as a page.
- Amazon RDS, Amazon MSK, Amazon OpenSearch Service, Amazon MQ, Amazon S3 — managed data services in the Platform account, consumed cross-account by workloads.
- GitHub Runners — CI/CD runners deployed as EKS workloads in the Platform account for centralized self-service pipelines.
- Network firewalls — security enforcement on the on-prem Direct Connect path.
Concepts & patterns¶
- concepts/control-plane-data-plane-separation — managed EKS control plane; AWS owns upgrades/patching/HA, teams own the data plane.
- concepts/tenant-isolation — multi-account + Cilium pod-level policy isolate application vs shared-services workloads.
- concepts/self-service-infrastructure — cloud-ops owns/governs the platform; app teams self-serve pre-configured namespaces + shared services.
- concepts/blast-radius — account boundaries + pod-level segmentation reduce the scope of potential incidents.
- concepts/observability — Hubble unifies previously fragmented, per-team network monitoring.
- Shared-services multi-account topology — the article's "North Star": Workloads / Platform / Network account split with a centrally-owned platform. Recorded as prose/tag here (single source); canonicalized into the existing concepts above rather than minted as a new page (AGENTS.md taxonomy gate).
Operational numbers¶
- 40% reduction in operational overhead (managed control plane eliminates self-managed upgrades/patching/HA).
- 4x increase in deployment frequency.
- 100% unified-architecture adoption across multiple business organizations.
- Topology: 3 AWS accounts (Workloads / Platform / Network), 2 AZs (us-west-1), 2 Regions (us-west-1 + us-east-2) with TGW inter-Region peering, Direct Connect Gateway with dual circuits to on-prem.
- Onboarding: weeks → days for new application teams.
Caveats¶
- Vendor/customer case study, not a production retrospective. The 40% / 4x / 100% figures are Equinix-reported outcomes with no methodology, baseline, or time window disclosed; treat them as directional, not benchmarked.
- The described diagram is the UAT environment, "with identical patterns replicated across non-production (system integration testing and development) and production" — production specifics are asserted, not shown.
- The post is written partly by Equinix authors (Manikandan Vasu, Vanji Sivajothy, Ramchandra Koty) and reads as an EKS advocacy narrative; the concrete architecture (account split, TGW hub, Cilium/Hubble, ALB+Gateway API vs NLB, GitHub Runners on EKS) is the durable content.
Source¶
- Original: https://aws.amazon.com/blogs/architecture/how-equinix-cut-operational-overhead-with-a-shared-services-architecture-on-amazon-eks/
- Raw markdown:
raw/aws/2026-09-17-how-equinix-cut-operational-overhead-with-a-shared-services-3a073a9f.md