Skip to content

AWS 2026-09-30 Tier 1

Read original ↗

Running multi-day AZ evacuation drills with ARC Zonal Shift

Summary

An AWS Architecture Blog walkthrough for running a sustained (48–72 hour) Availability Zone evacuation drill using Amazon Application Recovery Controller (ARC) Zonal Shift. The thesis: there is a gap between deploying multi-AZ and proving it works under sustained stress. Traditional DR tests shift traffic, confirm targets respond, and roll back within minutes — they validate the failover mechanism but never surface the time-dependent failure modes that only appear over hours or days. By shifting all traffic away from one AZ for days, you force a multi-tier platform (ALB+ECS Fargate, NLB+EKS, RDS for PostgreSQL, and Aurora PostgreSQL) to run full production load on N-1 zones, proving capacity sufficiency, database stability, client reconnection behaviour, and — most importantly — that teams can operate normally for days on reduced capacity. The post is explicitly aimed at financial-services operational-resilience programs that must produce auditable evidence of recovery ("show us the evidence", not "show us the runbook").

Key takeaways

  1. Short DR tests miss time-dependent failure modes; multi-day drills force them to play out. A multi-day N-1 run exposes: Auto Scaling policies not tuned for sustained N-1; deployment pipelines that don't validate AZ health before placing workloads; stale DNS / cached DB endpoints; time-based routine ops run against an N-1 architecture (cert/credential/secret rotations, maintenance windows, log rotation, backup automation, cron jobs); long-lived DB connections pinned to a specific AZ and used infrequently; and that recovery after a multi-day shift is a different operational procedure than a minutes-long rollback to a warm, near-identical AZ. This is a chaos-engineering/game-day discipline applied at the AZ-evacuation granularity.

  2. ARC Zonal Shift evacuates an AZ across all three traffic dimensions without touching application code. On a zonal shift, ARC takes two coordinated actions for Route 53 + Elastic Load Balancing: (1) DNS removal — the load balancer's IP in the affected AZ is removed from DNS so new client queries don't resolve to it; (2) cross-zone traffic blocking — LB nodes in the remaining AZs stop routing to targets in the shifted AZ even when cross-zone load balancing is enabled. For EKS clusters with zonal shift enabled, ARC goes further: cordons all nodes in the impacted AZ, removes pod endpoints in that AZ from EndpointSlice resources (redirecting east-west traffic), suspends AZ rebalancing for managed node groups and updates ASGs to launch only in healthy AZs, and preserves nodes/pods in the shifted AZ (not terminated) so full capacity is instantly available on restore. Combined with service-specific procedures this covers north-south ingress, east-west service-to-service, and outbound database connections. (Source: this article)

  3. Zonal Shift is a data-plane operation by design; the rest of the drill is control-plane. Because the shift works independently of the AWS control plane, it stays available even during an AZ impairment. The other steps (ECS service updates, manual RDS/Aurora failovers, subnet-group modifications) are control-plane operations. For a planned drill the control plane is healthy so this is moot; during a real impairment, prioritize the data-plane action (start the shift first to stop traffic immediately) and do control-plane work only after traffic is already shifted. This is a direct application of control-plane / data-plane separation. (Source: this article)

  4. A multi-day shift changes the operational surface area, not the mechanics. Starting a shift is identical for 1 hour or 72 hours; what changes over days: expiry management (shifts have a max duration — you must monitor and extend before expiry with update-zonal-shift or traffic silently returns); scaling drift (Auto Scaling in healthy AZs can create capacity imbalances — monitor and cap so recovery doesn't overload the returning AZ); connection-pool cycling (after 24h+ most client connections have recycled, truly validating DNS-TTL compliance across the whole client fleet — see connection pools); operational confidence (teams deploy/patch/troubleshoot in a reduced AZ environment); and safe recovery (restoring requires careful ordering — verify health, scale back gradually, reintroduce traffic incrementally). (Source: this article)

  5. Pre-scale for N-1 before the drill; don't scale reactively — this is static stability. "Your architecture should tolerate AZ loss without needing to scale reactively." In a 3-AZ environment, N-1 pre-scaling means running ~50% more compute than baseline peak. If that cost isn't justifiable, use scheduled scaling for drill windows, aggressive scale-out thresholds, or load shedding — but in an unplanned impairment there's no time to scale reactively, so unscaled workloads run degraded until scaling catches up (minutes under load). The article cites the Amazon Builders' Library static stability using Availability Zones. (Source: this article)

  6. The ELB/ECS tuning that makes a shift clean is in the connection-drain timings. Prerequisites: set ALB/NLB deregistration delay to 60s (vs default 300s) so connections drain quickly after a shift; configure target_group_health.dns_failover.minimum_healthy_targets.count on each target group; set ECS stopTimeout to 55s (just below the 60s dereg delay) so tasks finish in-flight requests before force-stop, avoiding 502s in the drain window. Zonal shift refuses single-AZ target groups — the ALB rejects the shift if healthy targets exist in only one AZ, so verify ≥2 AZs have registered targets first. (Source: this article)

  7. Database failover differs by engine; Aurora decouples storage from compute. For RDS Multi-AZ, if the primary is in the evacuated AZ you force failover with reboot-db-instance --force-failover; RDS recreates the standby in the evacuated AZ automatically (acceptable for a drill — it takes no client traffic); optionally relocate it via DB-subnet-group modification (snapshot → disable Multi-AZ → modify subnet group → re-enable, several minutes). For Aurora, storage is synchronously replicated to six storage nodes across AZs independently of compute, so only the writer/reader instance needs failover (failover-db-cluster --target-db-instance-identifier to a reader in a healthy AZ) — storage stays fully available. A pre-provisioned reader cuts failover from "<10 min" (promotion that must create an instance) to "<30s". (Source: this article)

  8. Outbound DB traffic from EKS pods isn't controlled by zonal shift. ARC zonal shift doesn't control pod→external-dependency connections. For AZ-specific DB endpoints, use Istio locality-aware routing (ServiceEntry + DestinationRule with localityLbSetting). The Aurora cluster endpoint auto-routes to the current writer regardless of AZ (no Istio needed for writer traffic); only AZ-specific reader endpoints need it. Verify pods use topologySpreadConstraints with maxSkew: 1 on topology.kubernetes.io/zone and are pre-scaled for N-1 — the shift doesn't evict pods or trigger autoscaling by itself. (Source: this article)

  9. The drill's value is the evidence package; monitor per-AZ throughout. A 48–72h evacuation requires continuous automated observation, not a visual confirmation. Keep per-AZ CloudWatch dashboards for each layer and export snapshots before/during/after; combine with the ARC zonal-shift event history (list-zonal-shifts) for an auditable evidence package. Metrics to watch per layer: ALB/NLB HealthyHostCount (→0 in evacuated AZ), RequestCount (zero in shifted AZ), TargetResponseTime and HTTPCode_Target_5XX_Count (capacity pressure in healthy AZs); ECS CPUUtilization <70–80% sustained; EKS node_status_condition (SchedulingDisabled in evacuated AZ); RDS DatabaseConnections (spike after failover → watch pool exhaustion) and ReplicaLag; Aurora AuroraReplicaLag (baseline ~100ms or less), CommitLatency, BufferCacheHitRatio (<99% = working set doesn't fit), VolumeBytesUsed (AZ-independent, should be unaffected). (Source: this article)

  10. Restore ordering matters. Verify the evacuated AZ is healthy → cancel the EKS zonal shift first (east-west resumes) → cancel the load balancer shift (north-south resumes as DNS propagates) → scale back added capacity gradually over 15–30 min after traffic redistributes → if cross-zone LB is disabled, confirm minimum_healthy_targets.count so Route 53 only sends traffic once the AZ has enough healthy targets → monitor per-AZ for 30 min. Zonal shift and zonal autoshift incur no additional charge. Progression advice: non-prod → prod low-traffic windows → sustained; then enable zonal autoshift so AWS shifts automatically on detected impairment, and use AWS Resilience Hub to assess RTO/RPO posture before/after. (Source: this article)

Architecture (illustrative multi-tier platform)

Layer Components Multi-AZ configuration
Traffic ingress ALB fronting ECS 3 AZs, cross-zone LB on
Compute (containers) ECS Fargate Stateless tasks across 3 AZ subnets
Traffic ingress NLB fronting EKS 3 AZs, cross-zone LB on
Compute (Kubernetes) EKS / EKS Auto Mode Stateless services, topology spread across 3 AZs
Database RDS for PostgreSQL Multi-AZ: primary AZ A, standby AZ B
Database Aurora PostgreSQL Writer AZ A, reader AZ B; storage replicated across all 3 AZs

The drill evacuates AZ A (the zone hosting the RDS primary and Aurora writer).

Key CLI commands

# Enable zonal shift on the LB (disabled by default)
aws elbv2 modify-load-balancer-attributes --load-balancer-arn $ALB_ARN \
  --attributes Key=zonal_shift.config.enabled,Value=true

# Start a 72h shift away from an AZ
aws arc-zonal-shift start-zonal-shift --resource-identifier $ALB_ARN \
  --away-from $AZ_ID_TO_EVACUATE --expires-in "72h" --comment "Multi-day AZ evacuation drill"

# Extend before expiry (else traffic returns automatically)
aws arc-zonal-shift update-zonal-shift --zonal-shift-id $SHIFT_ID \
  --resource-identifier $RESOURCE_ARN --expires-in "24h" --comment "Extending drill"

# One-time: enable zonal shift on an EKS cluster
aws eks update-cluster-config --name $CLUSTER_NAME --zonal-shift-config enabled=true

# RDS: force failover if primary is in the evacuated AZ
aws rds reboot-db-instance --db-instance-identifier $RDS_INSTANCE --force-failover

# Aurora: fail over writer to a reader in a healthy AZ
aws rds failover-db-cluster --db-cluster-identifier $CLUSTER_ID \
  --target-db-instance-identifier $READER_IN_HEALTHY_AZ

# Restore: cancel the shift(s)
aws arc-zonal-shift cancel-zonal-shift --zonal-shift-id $SHIFT_ID --resource-identifier $RESOURCE_ARN

Operational numbers / caveats

  • Drill duration: 48–72 hours; --expires-in has a maximum (monitor + extend).
  • N-1 pre-scaling in a 3-AZ env ≈ +50% compute over baseline peak.
  • ALB/NLB deregistration delay: 60s (vs default 300s); ECS stopTimeout: 55s.
  • Aurora storage: 6 nodes across AZs; pre-provisioned reader failover <30s vs <10 min without.
  • Aurora replica lag baseline: ~100ms or less; BufferCacheHitRatio alarm threshold: <99%.
  • ECS CPU sustained target: <70–80%. Restore scale-back window: 15–30 min; post-restore watch: 30 min.
  • Zonal shift and zonal autoshift are free (no additional charge, no ongoing resources created).
  • This is a reference/walkthrough post with illustrative architecture and planning numbers — not a production retrospective with measured incident data. The financial-services framing is motivational (regulatory operational-resilience trends), not a named customer case.

Systems / concepts extracted

Source

Last updated · 766 distilled / 2,225 read