Skip to content

AWS 2026-09-18

Read original ↗

ReadyOn's Four Walls of tenant isolation on Amazon EKS

Summary

ReadyOn, an AWS Partner running the Harmony workforce-intelligence platform for Fortune 100 enterprises, describes a multi-tenant Amazon EKS architecture that places four independent isolation layers between tenants — the "Four Walls" model. The core thesis: a Kubernetes namespace was never designed to be a security boundary, so a serious multi-tenant platform must layer isolation at independent abstraction layers so that crossing a tenant boundary requires simultaneously defeating the Kubernetes API, the node scheduler, the AWS software-defined network, and the data layer. The four walls are (1) namespace isolation (RBAC + admission control + Argo CD GitOps templating), (2) compute isolation (Karpenter dedicated node pools with a dual-taint strategy), (3) network isolation (per-tenant VPC security groups + default-deny network policies + private API endpoint), and (4) data isolation (per-tenant Aurora clusters, per-tenant KMS keys, and IRSA short-lived credentials). ReadyOn maps its defenses onto a 10-stage multi-tenant threat model keyed to MITRE ATT&CK, and validates that all six cross-tenant paths at the lateral-movement stage (Stage 8) are blocked.

This is a reference architecture for defense in depth applied to Kubernetes multi-tenancy: the layers are explicitly not claimed to be fully independent (they share admission control, GitOps, the EKS control plane, and IAM), so ReadyOn treats them as overlapping controls rather than independent probabilities.

Key takeaways

  1. Namespaces are not a security boundary. ReadyOn is explicit that Kubernetes designers never intended namespaces as a security boundary, and that a default Kubernetes deployment is not secure out of the box — historically anonymous requests reach the API server (denied only by RBAC), pods run as root by default, and network policies don't exist until someone writes them. The namespace is the logical foundation the other walls build on, not the isolation guarantee. (Source: sources/2026-09-18-aws-readyons-four-walls-of-tenant-isolation-on-amazon-eks)

  2. Multi-tenancy changes the threat model. For single-tenant clusters the question is "can an unauthorized user reach the cluster?" For multi-tenant clusters it becomes "can Tenant A access Tenant B's data?" — cross-tenant access, not perimeter breach, is the primary risk, with regulatory obligations across jurisdictions on the line.

  3. Wall 1 — Namespace isolation via GitOps templating. Every tenant namespace is generated from a single Git config entry through an Argo CD ApplicationSet, producing an identical security posture by construction — no "legacy" tenants with weaker policies. Per-namespace RoleBindings scope a tenant group to its own namespace; an admission-control framework enforces security contexts, blocks privileged configs, and forbids tenant creation of DaemonSets, ClusterRoles, and admission webhooks. The GitOps controller continuously reconciles live state against Git, reverting any manual/unauthorized change within seconds — a human operator never runs kubectl apply against production.

  4. Wall 2 — Compute isolation via Karpenter dual-taint. Karpenter provisioners create per-tenant nodes carrying two taints: a tenant-identifier taint (e.g. tenant=acme-corp) and a workload-type taint (e.g. workload=frontend). A pod must tolerate both to be scheduled, so Tenant A's frontend nodes are distinct from Tenant A's batch nodes and both are separate from Tenant B's. An admission controller validates that toleration claims match the namespace's tenant identity — even forged tolerations can't force cross-tenant placement. Nodes enforce IMDSv2 with a reduced hop limit (pods can't reach instance metadata), and each node's IAM role is scoped to a single tenant, limiting container-escape blast radius. (Source: sources/2026-09-18-aws-readyons-four-walls-of-tenant-isolation-on-amazon-eks)

  5. Wall 3 — Network isolation below Kubernetes. Each tenant's Aurora cluster sits behind a VPC security group that allows inbound only from that tenant's application-node security group — enforced by the AWS software-defined network, which a pod with limited access cannot manipulate. The VPC is split into four tiers: public perimeter (load balancers only), application (EKS worker nodes in private subnets), database (Aurora with no internet access), and control plane (private EKS API endpoint, VPN-only with OIDC + MFA). Kubernetes network policies enforce default-deny inter-namespace traffic; all other traffic is denied and logged, with VPC Flow Logs capturing traffic for anomaly detection.

  6. Wall 4 — Data isolation via dedicated Aurora + per-tenant KMS + IRSA. Rather than a shared database with row-level filtering (where a single missing WHERE tenant_id = ? leaks data), ReadyOn gives each tenant a dedicated Aurora cluster — "there is no row-level filter to forget." Each tenant's data-at-rest uses a distinct KMS key (envelope encryption), so even raw-storage access can't decrypt another tenant's data. Workloads authenticate via IRSA — short-lived STS credentials issued through OIDC federation between the EKS cluster and IAM, scoped per-tenant, expiring within minutes; application workloads hold no long-lived access keys. Secrets live in Secrets Manager under tenant-scoped paths, resolved at deploy time by the External Secrets Operator (zero secrets in Git). Per-tenant observability via OpenTelemetry keeps even error-rate spikes invisible across tenants.

  7. Threat model mapped to MITRE ATT&CK; Stage 8 is the inflection point. ReadyOn maps defenses onto a 10-stage multi-tenant threat model. The critical stage is Stage 8 — lateral movement, where a compromised Tenant A pod attempts to reach Tenant B. All six cross-tenant paths terminate at a wall: scheduling on another tenant's nodes → dual-taint + admission validation (Wall 2); reaching another DB → per-tenant security groups (Wall 3); cross-namespace traffic → default-deny network policies (Wall 3); reading secrets → IRSA scoping + scoped secrets paths (Wall 4); querying monitoring → per-tenant observability (Wall 4); creating an ingress route into another namespace → policy forbids tenant ingress modification and the platform ingress re-validates tenant context on every request (Wall 1). ReadyOn runs regular adversarial exercises simulating a full-access Tenant A pod; to date no test has produced a successful cross-tenant path.

  8. Innermost defense — Pod Security Standards "restricted". Every container runs runAsNonRoot: true, allowPrivilegeEscalation: false, readOnlyRootFilesystem: true, capabilities.drop: [ALL], and seccompProfile: RuntimeDefault. Combined with IMDSv2 restrictions and IRSA, even a container escape yields no useful node-level credentials — a hardening posture aligned with attack-surface minimization.

  9. GitOps as the security control plane. Every change begins as a reviewed pull request with automated policy checks; the Git repo is the single source of truth and the cluster reflects Git. Zero secrets in Git means even repo compromise yields no usable credentials — a secure-by-construction posture where the safe path is the only path.

  10. Operational benefits. 10 threat-model stages defended (all mapped to MITRE ATT&CK); no long-lived credentials in application workloads; consistent security posture across all tenants via GitOps templating (no legacy exceptions); self-healing drift correction within seconds; and a multi-region active-passive architecture with automated failover through Amazon Route 53.

Systems / concepts / patterns extracted

Operational numbers / specifics

  • 4 independent isolation layers (walls); 2 taints per node (tenant + workload-type); 10 threat-model stages; 6 cross-tenant attack paths blocked at Stage 8; 0 successful cross-tenant paths in adversarial testing to date.
  • IRSA credentials expire within minutes; drift reverted within seconds.
  • VPC split into 4 tiers (public / application / database / control plane).
  • Per-tenant KMS key → one tenant's key cannot decrypt another's data-at-rest.

Caveats

  • Layers are not fully independent. ReadyOn explicitly notes that admission control, GitOps, the EKS control plane, and IAM are shared control planes — hence "defense in depth" (overlapping controls) rather than a product of independent failure probabilities. A control-plane-level compromise (e.g. the admission webhook or the GitOps controller) could weaken multiple walls at once.
  • Cost / operational tradeoff. Dedicated per-tenant Aurora clusters, node pools, KMS keys, and observability instances trade the economics of shared infrastructure for isolation strength; the post frames this as acceptable for "some of the most sensitive data in enterprise IT" but it is not free.
  • This is a partner architecture narrative, not an AWS-service internals post — numbers are ReadyOn's own and there are no latency/throughput benchmarks.

Source

Last updated · 766 distilled / 2,225 read