Validating multi-Region DR for Terraform Enterprise with AWS FIS¶
Summary¶
After an October 2025 AWS regional service event in us-east-1 left Athenahealth's single-Region HashiCorp Terraform Enterprise (TFE) deployment inaccessible — blocking developers from deploying, modifying, or recovering any infrastructure — AWS, HashiCorp, and Athenahealth co-designed a customer-operated active-passive multi-Region DR architecture for TFE (us-east-1 primary, us-west-2 DR) and then validated it with AWS Fault Injection Service (FIS) rather than trusting the design on paper. The result: a 12–14 minute RTO and sub-1-minute RPO at ~30–40% higher cost than single-Region. The post is as much about the validation method — three progressive FIS phases (EC2/ASG, Aurora failover, S3 connectivity) plus a combined primary-Region-impairment run — as about the architecture. The headline lesson is a circular dependency: the failover scripts read infrastructure identifiers from Terraform S3 state files in the Region being recovered, so an S3 outage broke the very automation meant to escape it. Because TFE is officially single-Region only, the entire multi-Region failover is designed, operated, and tested by the customer — not a supported product configuration.
Key takeaways¶
-
Give your IaC control plane the same resilience as production. The founding insight: when TFE itself is unavailable during a regional event, you can't deploy fixes or recover infrastructure — the tool you'd reach for to recover is the tool that's down. This is why a mundane-seeming internal system got a full multi-Region DR treatment. (Source: this article)
-
The architecture is a textbook pilot light. Data (Aurora PostgreSQL global database, S3 state files) replicates continuously to us-west-2; compute stays at zero — the DR Auto Scaling group runs at minimum capacity 0 during normal operations. A warm standby variant (DR min capacity 1) trades higher cost for faster RTO. Maps directly onto the DR ladder. (Source: this article)
-
Four-step failover, sequenced to avoid split-brain. (1) Activate DR ASG 0→1 and pass health checks (2–5 min); (2) promote the Aurora global database writer (~1 min) — before DNS shifts, to prevent both Regions accepting writes (split-brain); (3) confirm Route 53 health-check-based DNS failover (~60 s, 60 s TTL on the ELB alias); (4) scale out for production load (5–10 min). Total execution: 12–14 min. The promotion-before-traffic ordering is enforced procedurally by the runbook, not by an automated control — the zero-capacity DR endpoint physically can't pass health checks until an operator runs the runbook. (Source: this article)
-
Failover must not depend on the Region's control plane — static stability. Route 53 record modification at failover time is a documented anti-pattern (the Route 53 control plane runs from a single Region); Athenahealth used pre-configured health-check-based failover routing that operates in the Route 53 data plane with no record edits during the event. Failover scripts run from outside the primary Region (DR Region, a management Region, or a CI/CD system) so they stay reachable. (Source: this article)
-
AWS FIS exposed real bugs that manual review missed. Phase A (
aws:ec2:stop-instances,aws:ec2:asg-insufficient-instance-capacity-error, ASG suspend/resume) confirmed 2–3 min ASG replacement — and surfaced an outdated AMI reference in the DR launch template (config drift that only appears when the ASG actually launches new instances). Phase B (aws:rds:failover-db-cluster) confirmed a 58-second Aurora writer promotion (~15 s write unavailability) — but revealed TFE connection pooling caused extended reconnection delays; they cut the pool timeout 60 s → 10 s (connection-pool tuning). (Source: this article) -
You can't disrupt S3 directly in FIS — disrupt the subnet. FIS has no direct S3-disruption action, so Phase C used
aws:network:disrupt-connectivitytargeting the TFE compute private subnets, injecting network-ACL rules that block egress to S3 endpoints (assumes an in-VPC path to S3, e.g. a gateway VPC endpoint). This tests TFE-as-S3-consumer, not S3 cross-Region replication itself (service-side replication doesn't traverse subnet NACLs). TFE detected loss in ~5 s; DR bucket held state with <30 s replication lag. (Source: this article) -
The circular-dependency pitfall (the marquee lesson). The failover/failback scripts fetched infrastructure identifiers (RDS global cluster ID, DR ASG name) via
terraform outputfrom state files stored in the primary-Region S3 bucket. During a regional impairment that bucket is unreachable, so the failover script hangs on an S3 timeout and stops — "you can't run the failover without access to the infrastructure you're trying to recover from." Phase C confirmed the severity. Fix: remove every recovery dependency on the Region you are recovering from — Athenahealth hardcoded identifiers directly in the scripts (with a CI/CD drift-detection job comparing them against Terraform outputs); a Region-independent config source (SSM Parameter Store replicated cross-Region, a DynamoDB global table, or values committed to the failover-scripts repo) achieves the same with less drift risk. (Source: this article) -
Test failback, not just failover. The state-file dependency broke both directions. Athenahealth validated failback by failing fully to DR, operating there 30 min, then returning — bidirectional S3 cross-Region replication prevented state-file loss and enabled failback without data resynchronization. Failback RTO ~20 min (incl. ~5 min to re-establish the Aurora global database). (Source: this article)
Architecture components¶
- Compute — EC2 running TFE application servers behind Auto Scaling groups across 3 AZs per Region; DR ASG scaled to 0 in steady state (scale-to-zero). HashiCorp's Terraform Enterprise Validated Design (HVD) module underpins the deployment (an HVD targets single-Region).
- Database — Aurora PostgreSQL-Compatible global database; primary cluster = 1 writer + 2 readers across 3 AZs; secondary maintains an inactive writer ready for promotion. Sub-second replication lag, managed failover.
- State storage — S3 for Terraform workspace state files, with bidirectional cross-Region replication between primary and DR buckets (enables failback without resync).
- Traffic — Route 53 DNS alias records → ELB Network Load
Balancer; 60 s TTL; failover routing policy driven by a Route 53
health check probing TFE's
/_health_check(200 OK) from a globally distributed checker fleet (detection independent of the primary Region). - Secrets/keys — Secrets Manager + KMS replicated cross-Region. Two TFE-specific gotchas: (a) the TFE encryption password protects the internal Vault unseal key and root token — DR instances configured with a different value can't start or decrypt existing data, so it must be replicated and referenced by the DR launch config; (b) in Active/Active TFE mode, external Redis holds the job queue and cache — plan a DR Redis equivalent and decide acceptable in-flight job loss.
- Observability/alerting — per-Region CloudWatch alarms on TFE instances, ASGs, Aurora clusters; alerts via SNS in both Regions.
- Future — Athenahealth plans to evaluate Amazon Application Recovery Controller (ARC) Region switch for an orchestrated, data-plane, explicitly-sequenced Region switch.
Operational numbers¶
- RTO (failover execution): 12–14 min (operator-triggered 4-step runbook).
- RPO: < 1 min (sub-second Aurora lag; up to 30 s S3 lag).
- EC2 failure recovery: 2–3 min (automated ASG replacement).
- Aurora failover: 1–2 min managed promotion (58 s measured; ~15 s write unavailability).
- Aurora replication lag: < 1 s (p99).
- S3 replication lag: < 30 s (p99).
- Data loss during failover: 0 bytes across each experiment (controlled conditions).
- TFE S3-loss detection: ~5 s. NLB health-check removal of failed instances: ~30 s.
- Connection-pool timeout: cut 60 s → 10 s.
- Failback RTO: ~20 min (incl. ~5 min Aurora global DB re-establishment).
- Cost: ~30–40% higher than single-Region (mostly Aurora global DB + S3 CRR).
- Engagement: 5-month cross-functional AWS + HashiCorp + Athenahealth effort.
Caveats¶
- Not a supported product configuration. TFE is officially single-Region-only and the HVD module targets single-Region; the multi-Region failover is entirely customer-designed/operated/tested.
- Measurements are controlled-test results. Both Aurora global DB and S3 CRR are asynchronous — an actual event can lose writes committed within the replication-lag window (sub-second Aurora, up to 30 s S3). Plan for near-zero, not zero RPO.
- Coordinated vs unplanned Aurora path differ. The measured 58 s used
Aurora's coordinated failover, which requires the primary Region
reachable to synchronize before promotion. A real primary-Region
impairment needs
failover-global-cluster --allow-data-loss(managed failover) or manual detach-and-promote — these don't wait for sync, so timing differs and RPO is bounded by lag at event time, not zero. - 12–14 min is execution time only. End-to-end recovery from event onset also includes detection + the decision to fail over — plan a larger overall RTO.
- Multi-Region DR isn't for every TFE deployment. Justified here because TFE manages infra for critical healthcare workloads; for less critical workloads a single-Region deployment with tested backup/restore may suffice.
Source¶
- Original: https://aws.amazon.com/blogs/architecture/validating-multi-region-dr-for-terraform-enterprise-with-aws-fis/
- Raw markdown:
raw/aws/2026-09-09-validating-multi-region-dr-for-terraform-enterprise-with-aws-7f2c1fc6.md
Related¶
- systems/terraform-enterprise — the IaC control plane being made multi-Region.
- systems/aws-fault-injection-service — the validation engine (three progressive phases + combined run).
- systems/aurora-global-database — cross-Region DB replication + managed failover.
- patterns/pilot-light-deployment — the DR tier this architecture instantiates.
- concepts/circular-dependency — the state-file-dependency pitfall.
- concepts/static-stability — failover must not depend on the recovering Region's control plane.
- concepts/rpo-rto — the budget dimensions met (12–14 min / <1 min).