Skip to content

SYSTEM Cited by 2 sources

Amazon Application Recovery Controller (ARC)

Definition

Amazon Application Recovery Controller (ARC) is an AWS service for application recovery readiness and automated failover across Regions and Availability Zones. It provides:

  • Readiness checks — validates that recovery configurations (cross-region replicas, DNS routing, capacity) are correctly provisioned and ready.
  • Routing controls — enables or disables traffic flow to specific cells/regions via highly-available data plane (cluster of 5 regional endpoints).
  • Safety rules — prevents accidental failover of too many cells simultaneously.

ARC sits at the orchestration layer of multi-region architectures — it doesn't replicate data or reroute traffic directly, but coordinates the signals that trigger failover actions.

ARC has two altitudes of recovery control:

  • Regional — readiness checks + routing controls + safety rules for Region-to-Region failover (above).
  • Zonal — Zonal Shift / zonal autoshift: evacuate a single Availability Zone within a Region (below). This is the more commonly exercised altitude because an AZ impairment is far more frequent than a Region loss.

Zonal Shift

Zonal Shift moves application traffic away from a single impaired Availability Zone without any application-code change, working natively across ALB, NLB, EKS clusters, and EC2 Auto Scaling groups. It is the mechanism behind both reactive AZ evacuation during a real impairment and proactive AZ evacuation drills (shifting for 48–72h to prove N-1 operation).

What ARC does on a shift (Route 53 + ELB)

When you start a zonal shift, ARC takes two coordinated actions:

  1. DNS removal — the load balancer's IP address in the affected AZ is removed from DNS, so new client queries no longer resolve to that endpoint.
  2. Cross-zone traffic blocking — load-balancer nodes in the remaining AZs stop routing requests to targets in the shifted AZ, even when cross-zone load balancing is enabled (the ALB default). Targets in the shifted AZ are fully isolated regardless of cross-zone configuration.

Zonal shift refuses single-AZ target groups: the ALB rejects the shift if healthy targets exist in only one AZ. Verify each target group has registered targets in ≥2 AZs before shifting. (Source: sources/2026-09-30-aws-running-multi-day-az-evacuation-drills-with-arc-zonal-shift)

What ARC additionally does for EKS

For EKS clusters with zonal shift enabled, ARC handles both the infrastructure and the Kubernetes networking layer automatically:

  • Cordons all nodes in the impacted AZ (no new pod scheduling).
  • Removes pod endpoints in the impacted AZ from EndpointSlice resources, so east-west (service-to-service) traffic is redirected to pods in healthy AZs by the built-in Kubernetes EndpointSlice controller.
  • Suspends AZ rebalancing for managed node groups and updates ASGs to launch instances only in healthy AZs.
  • Preserves nodes and pods in the shifted AZ — they are not terminated — so full capacity is immediately available when the shift ends.

Combined with service-specific procedures (ECS task redistribution, RDS/Aurora failover) this produces a complete AZ evacuation across all three traffic dimensions: north-south ingress, east-west service communication, and outbound database connections. ARC zonal shift does not control pod→external (e.g. AZ-specific RDS endpoint) connections — use Istio locality-aware routing for those. (Source: sources/2026-09-30-aws-running-multi-day-az-evacuation-drills-with-arc-zonal-shift)

Data plane by design

Zonal Shift is a data-plane operation — it works independently of the AWS control plane and therefore remains available even during an AZ impairment. This is a defining control-plane / data-plane separation property: during a real impairment you can always start the shift to stop traffic immediately, whereas the surrounding steps (ECS service updates, manual RDS/Aurora failover, subnet-group edits) are control-plane operations that may be degraded. Operational rule: start the zonal shift first, do control-plane work after traffic has already shifted. (Source: sources/2026-09-30-aws-running-multi-day-az-evacuation-drills-with-arc-zonal-shift)

Expiry management (multi-day shifts)

Shifts have a maximum duration set via --expires-in; if a shift expires before you cancel it, traffic automatically returns to the shifted AZ. For a multi-day drill you must monitor remaining time and extend before expiry with update-zonal-shift. The mechanics of starting a shift are identical for 1 hour or 72 hours — what changes over days is the operational surface area (scaling drift in healthy AZs, connection-pool cycling, and careful ordered recovery). (Source: sources/2026-09-30-aws-running-multi-day-az-evacuation-drills-with-arc-zonal-shift)

Core CLI

# Enable zonal shift on the LB (disabled by default)
aws elbv2 modify-load-balancer-attributes --load-balancer-arn $ALB_ARN \
  --attributes Key=zonal_shift.config.enabled,Value=true

aws arc-zonal-shift start-zonal-shift  --resource-identifier $ARN --away-from $AZ_ID --expires-in "72h"
aws arc-zonal-shift update-zonal-shift --zonal-shift-id $ID --resource-identifier $ARN --expires-in "24h"
aws arc-zonal-shift cancel-zonal-shift --zonal-shift-id $ID --resource-identifier $ARN
aws eks update-cluster-config --name $CLUSTER --zonal-shift-config enabled=true

Zonal shift and zonal autoshift incur no additional charge and create no ongoing resources. The ARC zonal-shift event history (list-zonal-shifts) is a primary input to an auditable drill-evidence package.

Zonal autoshift

Zonal autoshift lets AWS shift traffic away from an AZ automatically when internal telemetry detects a potential impairment — the fully-automated evolution of a manual shift, recommended once teams have built confidence through drills. AWS FIS has a dedicated recovery action for zonal autoshift, so autoshift behaviour can itself be chaos-tested.

Relationship to the broader resilience toolchain

  • AWS FIS — injects AZ impairment (the AZ Availability: Power Interruption scenario) to validate that an AZ evacuation / autoshift actually works; ARC is the response mechanism.
  • AWS Resilience Hub — assesses RTO/RPO posture before and after an evacuation drill.
  • concepts/static-stability — the drill only works if the fleet is pre-scaled for N-1 so a shifted AZ's load is absorbed without reactive scaling.

Pattern of appearance

ARC typically appears as the automated failover trigger in active-active or warm-standby architectures. CloudWatch alarms detect degradation → ARC routing controls shift traffic → downstream systems (Route 53, ALB, custom handlers) respond.

Seen in

Last updated · 766 distilled / 2,225 read