Running multi-day AZ evacuation drills with ARC Zonal Shift¶
Summary¶
An AWS Architecture Blog walkthrough for running a sustained (48–72 hour) Availability Zone evacuation drill using Amazon Application Recovery Controller (ARC) Zonal Shift. The thesis: there is a gap between deploying multi-AZ and proving it works under sustained stress. Traditional DR tests shift traffic, confirm targets respond, and roll back within minutes — they validate the failover mechanism but never surface the time-dependent failure modes that only appear over hours or days. By shifting all traffic away from one AZ for days, you force a multi-tier platform (ALB+ECS Fargate, NLB+EKS, RDS for PostgreSQL, and Aurora PostgreSQL) to run full production load on N-1 zones, proving capacity sufficiency, database stability, client reconnection behaviour, and — most importantly — that teams can operate normally for days on reduced capacity. The post is explicitly aimed at financial-services operational-resilience programs that must produce auditable evidence of recovery ("show us the evidence", not "show us the runbook").
Key takeaways¶
-
Short DR tests miss time-dependent failure modes; multi-day drills force them to play out. A multi-day N-1 run exposes: Auto Scaling policies not tuned for sustained N-1; deployment pipelines that don't validate AZ health before placing workloads; stale DNS / cached DB endpoints; time-based routine ops run against an N-1 architecture (cert/credential/secret rotations, maintenance windows, log rotation, backup automation, cron jobs); long-lived DB connections pinned to a specific AZ and used infrequently; and that recovery after a multi-day shift is a different operational procedure than a minutes-long rollback to a warm, near-identical AZ. This is a chaos-engineering/game-day discipline applied at the AZ-evacuation granularity.
-
ARC Zonal Shift evacuates an AZ across all three traffic dimensions without touching application code. On a zonal shift, ARC takes two coordinated actions for Route 53 + Elastic Load Balancing: (1) DNS removal — the load balancer's IP in the affected AZ is removed from DNS so new client queries don't resolve to it; (2) cross-zone traffic blocking — LB nodes in the remaining AZs stop routing to targets in the shifted AZ even when cross-zone load balancing is enabled. For EKS clusters with zonal shift enabled, ARC goes further: cordons all nodes in the impacted AZ, removes pod endpoints in that AZ from EndpointSlice resources (redirecting east-west traffic), suspends AZ rebalancing for managed node groups and updates ASGs to launch only in healthy AZs, and preserves nodes/pods in the shifted AZ (not terminated) so full capacity is instantly available on restore. Combined with service-specific procedures this covers north-south ingress, east-west service-to-service, and outbound database connections. (Source: this article)
-
Zonal Shift is a data-plane operation by design; the rest of the drill is control-plane. Because the shift works independently of the AWS control plane, it stays available even during an AZ impairment. The other steps (ECS service updates, manual RDS/Aurora failovers, subnet-group modifications) are control-plane operations. For a planned drill the control plane is healthy so this is moot; during a real impairment, prioritize the data-plane action (start the shift first to stop traffic immediately) and do control-plane work only after traffic is already shifted. This is a direct application of control-plane / data-plane separation. (Source: this article)
-
A multi-day shift changes the operational surface area, not the mechanics. Starting a shift is identical for 1 hour or 72 hours; what changes over days: expiry management (shifts have a max duration — you must monitor and extend before expiry with
update-zonal-shiftor traffic silently returns); scaling drift (Auto Scaling in healthy AZs can create capacity imbalances — monitor and cap so recovery doesn't overload the returning AZ); connection-pool cycling (after 24h+ most client connections have recycled, truly validating DNS-TTL compliance across the whole client fleet — see connection pools); operational confidence (teams deploy/patch/troubleshoot in a reduced AZ environment); and safe recovery (restoring requires careful ordering — verify health, scale back gradually, reintroduce traffic incrementally). (Source: this article) -
Pre-scale for N-1 before the drill; don't scale reactively — this is static stability. "Your architecture should tolerate AZ loss without needing to scale reactively." In a 3-AZ environment, N-1 pre-scaling means running ~50% more compute than baseline peak. If that cost isn't justifiable, use scheduled scaling for drill windows, aggressive scale-out thresholds, or load shedding — but in an unplanned impairment there's no time to scale reactively, so unscaled workloads run degraded until scaling catches up (minutes under load). The article cites the Amazon Builders' Library static stability using Availability Zones. (Source: this article)
-
The ELB/ECS tuning that makes a shift clean is in the connection-drain timings. Prerequisites: set ALB/NLB deregistration delay to 60s (vs default 300s) so connections drain quickly after a shift; configure
target_group_health.dns_failover.minimum_healthy_targets.counton each target group; set ECSstopTimeoutto 55s (just below the 60s dereg delay) so tasks finish in-flight requests before force-stop, avoiding 502s in the drain window. Zonal shift refuses single-AZ target groups — the ALB rejects the shift if healthy targets exist in only one AZ, so verify ≥2 AZs have registered targets first. (Source: this article) -
Database failover differs by engine; Aurora decouples storage from compute. For RDS Multi-AZ, if the primary is in the evacuated AZ you force failover with
reboot-db-instance --force-failover; RDS recreates the standby in the evacuated AZ automatically (acceptable for a drill — it takes no client traffic); optionally relocate it via DB-subnet-group modification (snapshot → disable Multi-AZ → modify subnet group → re-enable, several minutes). For Aurora, storage is synchronously replicated to six storage nodes across AZs independently of compute, so only the writer/reader instance needs failover (failover-db-cluster --target-db-instance-identifierto a reader in a healthy AZ) — storage stays fully available. A pre-provisioned reader cuts failover from "<10 min" (promotion that must create an instance) to "<30s". (Source: this article) -
Outbound DB traffic from EKS pods isn't controlled by zonal shift. ARC zonal shift doesn't control pod→external-dependency connections. For AZ-specific DB endpoints, use Istio locality-aware routing (
ServiceEntry+DestinationRulewithlocalityLbSetting). The Aurora cluster endpoint auto-routes to the current writer regardless of AZ (no Istio needed for writer traffic); only AZ-specific reader endpoints need it. Verify pods usetopologySpreadConstraintswithmaxSkew: 1ontopology.kubernetes.io/zoneand are pre-scaled for N-1 — the shift doesn't evict pods or trigger autoscaling by itself. (Source: this article) -
The drill's value is the evidence package; monitor per-AZ throughout. A 48–72h evacuation requires continuous automated observation, not a visual confirmation. Keep per-AZ CloudWatch dashboards for each layer and export snapshots before/during/after; combine with the ARC zonal-shift event history (
list-zonal-shifts) for an auditable evidence package. Metrics to watch per layer: ALB/NLBHealthyHostCount(→0 in evacuated AZ),RequestCount(zero in shifted AZ),TargetResponseTimeandHTTPCode_Target_5XX_Count(capacity pressure in healthy AZs); ECSCPUUtilization<70–80% sustained; EKSnode_status_condition(SchedulingDisabledin evacuated AZ); RDSDatabaseConnections(spike after failover → watch pool exhaustion) andReplicaLag; AuroraAuroraReplicaLag(baseline ~100ms or less),CommitLatency,BufferCacheHitRatio(<99% = working set doesn't fit),VolumeBytesUsed(AZ-independent, should be unaffected). (Source: this article) -
Restore ordering matters. Verify the evacuated AZ is healthy → cancel the EKS zonal shift first (east-west resumes) → cancel the load balancer shift (north-south resumes as DNS propagates) → scale back added capacity gradually over 15–30 min after traffic redistributes → if cross-zone LB is disabled, confirm
minimum_healthy_targets.countso Route 53 only sends traffic once the AZ has enough healthy targets → monitor per-AZ for 30 min. Zonal shift and zonal autoshift incur no additional charge. Progression advice: non-prod → prod low-traffic windows → sustained; then enable zonal autoshift so AWS shifts automatically on detected impairment, and use AWS Resilience Hub to assess RTO/RPO posture before/after. (Source: this article)
Architecture (illustrative multi-tier platform)¶
| Layer | Components | Multi-AZ configuration |
|---|---|---|
| Traffic ingress | ALB fronting ECS | 3 AZs, cross-zone LB on |
| Compute (containers) | ECS Fargate | Stateless tasks across 3 AZ subnets |
| Traffic ingress | NLB fronting EKS | 3 AZs, cross-zone LB on |
| Compute (Kubernetes) | EKS / EKS Auto Mode | Stateless services, topology spread across 3 AZs |
| Database | RDS for PostgreSQL | Multi-AZ: primary AZ A, standby AZ B |
| Database | Aurora PostgreSQL | Writer AZ A, reader AZ B; storage replicated across all 3 AZs |
The drill evacuates AZ A (the zone hosting the RDS primary and Aurora writer).
Key CLI commands¶
# Enable zonal shift on the LB (disabled by default)
aws elbv2 modify-load-balancer-attributes --load-balancer-arn $ALB_ARN \
--attributes Key=zonal_shift.config.enabled,Value=true
# Start a 72h shift away from an AZ
aws arc-zonal-shift start-zonal-shift --resource-identifier $ALB_ARN \
--away-from $AZ_ID_TO_EVACUATE --expires-in "72h" --comment "Multi-day AZ evacuation drill"
# Extend before expiry (else traffic returns automatically)
aws arc-zonal-shift update-zonal-shift --zonal-shift-id $SHIFT_ID \
--resource-identifier $RESOURCE_ARN --expires-in "24h" --comment "Extending drill"
# One-time: enable zonal shift on an EKS cluster
aws eks update-cluster-config --name $CLUSTER_NAME --zonal-shift-config enabled=true
# RDS: force failover if primary is in the evacuated AZ
aws rds reboot-db-instance --db-instance-identifier $RDS_INSTANCE --force-failover
# Aurora: fail over writer to a reader in a healthy AZ
aws rds failover-db-cluster --db-cluster-identifier $CLUSTER_ID \
--target-db-instance-identifier $READER_IN_HEALTHY_AZ
# Restore: cancel the shift(s)
aws arc-zonal-shift cancel-zonal-shift --zonal-shift-id $SHIFT_ID --resource-identifier $RESOURCE_ARN
Operational numbers / caveats¶
- Drill duration: 48–72 hours;
--expires-inhas a maximum (monitor + extend). - N-1 pre-scaling in a 3-AZ env ≈ +50% compute over baseline peak.
- ALB/NLB deregistration delay: 60s (vs default 300s); ECS
stopTimeout: 55s. - Aurora storage: 6 nodes across AZs; pre-provisioned reader failover <30s vs <10 min without.
- Aurora replica lag baseline: ~100ms or less; BufferCacheHitRatio alarm threshold: <99%.
- ECS CPU sustained target: <70–80%. Restore scale-back window: 15–30 min; post-restore watch: 30 min.
- Zonal shift and zonal autoshift are free (no additional charge, no ongoing resources created).
- This is a reference/walkthrough post with illustrative architecture and planning numbers — not a production retrospective with measured incident data. The financial-services framing is motivational (regulatory operational-resilience trends), not a named customer case.
Systems / concepts extracted¶
- Systems: ARC Zonal Shift (+ zonal autoshift), ECS/Fargate, EKS/EKS Auto Mode, RDS for PostgreSQL, Aurora PostgreSQL, ALB, NLB, Route 53, CloudWatch, AWS FIS (zonal-autoshift FIS recovery action), Istio (locality-aware routing), Karpenter/managed node groups, Resilience Hub.
- Concepts: static stability (N-1 pre-scaling, Builders' Library), chaos engineering (multi-day game-day), control-plane/data-plane separation (data-plane shift stays up during impairment), graceful degradation (unscaled workloads run degraded until scaling catches up), blast radius (single-AZ isolation), RTO/RPO (Resilience Hub validation), connection-pool cycling (24h+ validates DNS-TTL compliance), regional/zonal failover.
- Taxonomy note: the "multi-day AZ evacuation drill / game-day" idea is recorded as prose + tags and
mapped into concepts/chaos-engineering (which already covers whole-AZ-partition / progressive-scope),
not minted as a standalone pattern (single-source article-specific framing; "when unsure, tag — don't
mint"). Connection-drain / deregistration-delay tuning and
topologySpreadConstraints maxSkew:1are implementation detail, left as prose.
Source¶
- Original: https://aws.amazon.com/blogs/architecture/running-multi-day-az-evacuation-drills-with-arc-zonal-shift/
- Raw markdown:
raw/aws/2026-09-30-running-multi-day-az-evacuation-drills-with-arc-zonal-shift-50645052.md
Related¶
- systems/amazon-application-recovery-controller — the service; this source is its canonical Zonal Shift home.
- systems/aws-fault-injection-service — the chaos-injection sibling (AZ Power Interruption scenario; zonal-autoshift FIS recovery action).
- concepts/static-stability — N-1 pre-scaling is the drill's load-bearing prerequisite.
- concepts/chaos-engineering — the discipline this operationalizes at AZ-evacuation granularity.
- concepts/control-plane-data-plane-separation — why the shift stays available during a real impairment.
- concepts/graceful-degradation — the degraded state unscaled workloads fall into under an unplanned impairment.