CONCEPT Cited by 3 sources
Regional failover¶
What it is¶
Regional failover is the act of shifting a service's traffic away from a failed (or degraded) geographic region into one or more healthy regions, so that users continue to be served despite the loss of an entire region's capacity. It is the coarsest-grained rung of the reliability ladder — above node, rack, and availability-zone redundancy — and the one with the largest blast radius to absorb, because the surviving regions must suddenly carry load they were not sized for.
A regional failover is "in one sense a significant event," but a well-designed system engineers it to minimize customer impact — the goal is that most users never notice. (Source: sources/2026-09-16-spotify-ai-changed-how-spotify-builds-quality-at-higher-velocity)
The capacity problem¶
The defining hard part of regional failover is capacity in the receiving region. Failover only works if the surviving region(s) have enough spare compute to absorb the redirected load. Two things make this fragile:
- Headroom is expensive. Reserving idle capacity for a "once in a blue moon" regional loss (elasticity / headroom) costs money every day for an event that rarely fires — so it is chronically under-provisioned.
- Spare capacity is no longer guaranteed. Spotify observed that the industry-wide AI demand spike on CPUs and GPUs shrank the spare capacity it used to be able to count on. Failovers that were previously trivial and unnoticeable became user-visible once the receiving region couldn't reliably absorb the load. (Source: sources/2026-09-16-spotify-ai-changed-how-spotify-builds-quality-at-higher-velocity)
When you cannot guarantee capacity for everything, failover forces an explicit choice: shed the lower tiers of service to protect the critical path. Spotify now frames it as "when we failover, we must accept that there may not be capacity for the lower tiers of services" — i.e. regional failover degrades into graceful degradation under capacity constraint, where criticality-based service tiering decides who keeps capacity.
Doing it gradually¶
Dumping an entire region's traffic onto a neighbor instantaneously can overwhelm the receiving region and cascade the failure. The mitigations are:
- Move traffic gradually while confirming the receiving region can absorb the additional load and stay stable — a staged traffic shift rather than a hard cut.
- Reserve edge capacity. Spotify doubled reserved edge capacity after a May 2026 incident so the network edge can hold redirected traffic. (Source: sources/2026-09-16-spotify-ai-changed-how-spotify-builds-quality-at-higher-velocity)
- Test regional spillover explicitly, and build the control to shift traffic (Spotify can shift part of its internal service-mesh traffic manually today; extending that control to edge traffic and testing spillover were still in progress as of Sep 2026).
Failover topologies¶
Regional failover shows up in two broad shapes across the corpus:
- Active-passive — a standby region is promoted when the primary fails (common in disaster-recovery designs; pairs with async cross-region replication to keep the standby warm).
- Active-active — multiple regions serve simultaneously and a failure simply reweights traffic toward the survivors. This is what makes the capacity problem acute: the survivors were already serving their own load.
Seen in¶
-
sources/2026-09-30-aws-running-multi-day-az-evacuation-drills-with-arc-zonal-shift — zonal failover: the within-Region sibling of regional failover. Rather than shifting across Regions, ARC Zonal Shift shifts traffic away from a single AZ (N-1) — the same traffic-shifting + capacity-planning problem at a smaller blast radius. The post is notable for insisting the shift must be pre-scaled (static stability, ~+50% in a 3-AZ env) and for distinguishing a brief failover test (validates the mechanism) from a sustained 48–72h evacuation (validates operating under N-1 for days). Restore is explicitly incremental: cancel east-west (EKS) first, then north-south (LB), then scale back gradually over 15–30 min — don't return all traffic at once. (Source: sources/2026-09-30-aws-running-multi-day-az-evacuation-drills-with-arc-zonal-shift)
-
sources/2026-09-16-spotify-ai-changed-how-spotify-builds-quality-at-higher-velocity — Spotify shifts traffic to another region when a region fails; the 2026 compute shortage made previously-trivial failovers user-visible, forcing an edge/tiering rework: accept that lower service tiers may lose capacity on failover, double reserved edge capacity, move traffic gradually, and test regional spillover.
- sources/2026-09-17-aws-building-cloud-native-pacs-on-aws — the cloud-native PACS failure spectrum: a two-AZ cloud deployment with automatic NLB failover to the surviving AZ "within seconds"; if a local server fails, reads route to the cloud (recent data already synced); if the cloud link drops, the local cache continues serving recent studies uninterrupted — bidirectional local↔cloud fallback rather than one-way regional failover.
Related¶
- concepts/graceful-degradation — what a capacity-constrained failover degrades into: shed lower tiers to protect critical paths.
- concepts/elasticity — the headroom problem that makes receiving-region capacity fragile.
- concepts/blast-radius — a regional failover is the largest failure domain to absorb.
- concepts/backpressure — workload controls / prioritization that decide who keeps capacity under load.
- patterns/staged-rollout — move traffic gradually rather than a hard cut.
- patterns/fast-rollback — the adjacent "undo quickly" safeguard.
- patterns/async-replication-for-cross-region — keeps a standby region warm for active-passive failover.