SYSTEM Cited by 2 sources
Amazon Application Recovery Controller (ARC)¶
Definition¶
Amazon Application Recovery Controller (ARC) is an AWS service for application recovery readiness and automated failover across Regions and Availability Zones. It provides:
- Readiness checks — validates that recovery configurations (cross-region replicas, DNS routing, capacity) are correctly provisioned and ready.
- Routing controls — enables or disables traffic flow to specific cells/regions via highly-available data plane (cluster of 5 regional endpoints).
- Safety rules — prevents accidental failover of too many cells simultaneously.
ARC sits at the orchestration layer of multi-region architectures — it doesn't replicate data or reroute traffic directly, but coordinates the signals that trigger failover actions.
ARC has two altitudes of recovery control:
- Regional — readiness checks + routing controls + safety rules for Region-to-Region failover (above).
- Zonal — Zonal Shift / zonal autoshift: evacuate a single Availability Zone within a Region (below). This is the more commonly exercised altitude because an AZ impairment is far more frequent than a Region loss.
Zonal Shift¶
Zonal Shift moves application traffic away from a single impaired Availability Zone without any application-code change, working natively across ALB, NLB, EKS clusters, and EC2 Auto Scaling groups. It is the mechanism behind both reactive AZ evacuation during a real impairment and proactive AZ evacuation drills (shifting for 48–72h to prove N-1 operation).
What ARC does on a shift (Route 53 + ELB)¶
When you start a zonal shift, ARC takes two coordinated actions:
- DNS removal — the load balancer's IP address in the affected AZ is removed from DNS, so new client queries no longer resolve to that endpoint.
- Cross-zone traffic blocking — load-balancer nodes in the remaining AZs stop routing requests to targets in the shifted AZ, even when cross-zone load balancing is enabled (the ALB default). Targets in the shifted AZ are fully isolated regardless of cross-zone configuration.
Zonal shift refuses single-AZ target groups: the ALB rejects the shift if healthy targets exist in only one AZ. Verify each target group has registered targets in ≥2 AZs before shifting. (Source: sources/2026-09-30-aws-running-multi-day-az-evacuation-drills-with-arc-zonal-shift)
What ARC additionally does for EKS¶
For EKS clusters with zonal shift enabled, ARC handles both the infrastructure and the Kubernetes networking layer automatically:
- Cordons all nodes in the impacted AZ (no new pod scheduling).
- Removes pod endpoints in the impacted AZ from EndpointSlice resources, so east-west (service-to-service) traffic is redirected to pods in healthy AZs by the built-in Kubernetes EndpointSlice controller.
- Suspends AZ rebalancing for managed node groups and updates ASGs to launch instances only in healthy AZs.
- Preserves nodes and pods in the shifted AZ — they are not terminated — so full capacity is immediately available when the shift ends.
Combined with service-specific procedures (ECS task redistribution, RDS/Aurora failover) this produces a complete AZ evacuation across all three traffic dimensions: north-south ingress, east-west service communication, and outbound database connections. ARC zonal shift does not control pod→external (e.g. AZ-specific RDS endpoint) connections — use Istio locality-aware routing for those. (Source: sources/2026-09-30-aws-running-multi-day-az-evacuation-drills-with-arc-zonal-shift)
Data plane by design¶
Zonal Shift is a data-plane operation — it works independently of the AWS control plane and therefore remains available even during an AZ impairment. This is a defining control-plane / data-plane separation property: during a real impairment you can always start the shift to stop traffic immediately, whereas the surrounding steps (ECS service updates, manual RDS/Aurora failover, subnet-group edits) are control-plane operations that may be degraded. Operational rule: start the zonal shift first, do control-plane work after traffic has already shifted. (Source: sources/2026-09-30-aws-running-multi-day-az-evacuation-drills-with-arc-zonal-shift)
Expiry management (multi-day shifts)¶
Shifts have a maximum duration set via --expires-in; if a shift expires
before you cancel it, traffic automatically returns to the shifted AZ. For a
multi-day drill you must monitor remaining time and extend before expiry with
update-zonal-shift. The mechanics of starting a shift are identical for 1 hour
or 72 hours — what changes over days is the operational surface area (scaling
drift in healthy AZs, connection-pool cycling, and careful ordered recovery).
(Source: sources/2026-09-30-aws-running-multi-day-az-evacuation-drills-with-arc-zonal-shift)
Core CLI¶
# Enable zonal shift on the LB (disabled by default)
aws elbv2 modify-load-balancer-attributes --load-balancer-arn $ALB_ARN \
--attributes Key=zonal_shift.config.enabled,Value=true
aws arc-zonal-shift start-zonal-shift --resource-identifier $ARN --away-from $AZ_ID --expires-in "72h"
aws arc-zonal-shift update-zonal-shift --zonal-shift-id $ID --resource-identifier $ARN --expires-in "24h"
aws arc-zonal-shift cancel-zonal-shift --zonal-shift-id $ID --resource-identifier $ARN
aws eks update-cluster-config --name $CLUSTER --zonal-shift-config enabled=true
Zonal shift and zonal autoshift incur no additional charge and create no
ongoing resources. The ARC zonal-shift event history (list-zonal-shifts) is a
primary input to an auditable drill-evidence package.
Zonal autoshift¶
Zonal autoshift lets AWS shift traffic away from an AZ automatically when internal telemetry detects a potential impairment — the fully-automated evolution of a manual shift, recommended once teams have built confidence through drills. AWS FIS has a dedicated recovery action for zonal autoshift, so autoshift behaviour can itself be chaos-tested.
Relationship to the broader resilience toolchain¶
- AWS FIS — injects AZ impairment (the AZ Availability: Power Interruption scenario) to validate that an AZ evacuation / autoshift actually works; ARC is the response mechanism.
- AWS Resilience Hub — assesses RTO/RPO posture before and after an evacuation drill.
- concepts/static-stability — the drill only works if the fleet is pre-scaled for N-1 so a shifted AZ's load is absorbed without reactive scaling.
Pattern of appearance¶
ARC typically appears as the automated failover trigger in active-active or warm-standby architectures. CloudWatch alarms detect degradation → ARC routing controls shift traffic → downstream systems (Route 53, ALB, custom handlers) respond.
Seen in¶
-
sources/2026-09-30-aws-running-multi-day-az-evacuation-drills-with-arc-zonal-shift — canonical wiki home for Zonal Shift. Multi-day (48–72h) AZ evacuation drill across ALB+ECS Fargate, NLB+EKS, RDS and Aurora. Documents the two coordinated shift actions (DNS removal + cross-zone blocking), the EKS EndpointSlice/cordon/ASG behaviour, the data-plane-by-design property (shift stays up during a real impairment → "shift first, control-plane after"), expiry management for multi-day shifts, the single-AZ-target-group refusal, zonal autoshift, and the tie to N-1 pre-scaling and game-day discipline. Zonal shift is free (AWS Architecture Blog, 2026-09-30).
-
sources/2026-07-22-aws-building-multi-region-resiliency-for-cloudformation-custom-resources — ARC provides automated failover for CloudFormation custom resource processing; CloudWatch alarms on SQS depth and Lambda health feed into ARC-triggered region failover without manual intervention (AWS Architecture Blog, 2026-07-22).
Related¶
- systems/aws-fault-injection-service — injects AZ impairment (AZ Power Interruption scenario) to validate the response ARC provides; has a zonal-autoshift recovery action.
- systems/aws-resilience-hub — RTO/RPO posture assessment before/after a drill.
- systems/aws-application-load-balancer, systems/aws-nlb — the ELB resources a zonal shift operates on (DNS removal + cross-zone blocking).
- systems/aws-eks — the EKS cluster shift (EndpointSlice isolation, node cordoning).
- systems/amazon-route53 — the DNS layer ARC manipulates on a shift.
- concepts/static-stability — N-1 pre-scaling is the prerequisite that makes an evacuation absorb load without reactive scaling.
- concepts/control-plane-data-plane-separation — why Zonal Shift (data plane) stays available during an AZ impairment.
- concepts/chaos-engineering — the discipline AZ evacuation drills operationalize.