Skip to content

SYSTEM Cited by 1 source

Terraform Enterprise (TFE)

What it is

Terraform Enterprise (TFE) is HashiCorp's self-managed, on-premises distribution of Terraform — the paid, self-hosted counterpart to HCP Terraform (Terraform Cloud). It runs Terraform runs (plan/apply), stores workspace state files, manages workspaces, policy (Sentinel/OPA), and run pipelines inside a customer's own infrastructure. On AWS, HashiCorp ships a Terraform Enterprise Validated Design (HVD) module that stands up the reference deployment.

TFE is an infrastructure control plane: it is the tool teams use to deploy, modify, and recover other infrastructure. That framing is what makes its own availability a resilience concern — if TFE is down during an incident, you can't deploy the fix or recover the infrastructure it manages. (Source: sources/2026-09-09-aws-validating-multi-region-dr-for-terraform-enterprise-with-aws-fis)

Reference deployment on AWS (HVD)

The AWS-validated TFE deployment comprises:

  • Compute — EC2 running TFE application servers, fronted by an Auto Scaling group across 3 AZs.
  • Application state — Aurora PostgreSQL-Compatible for TFE's relational state.
  • Workspace state files — S3 stores Terraform workspace terraform.tfstate files.

The HVD module and HashiCorp's supported configuration target a single AWS Region. AZ-level resilience comes for free; Regional failover does not — it must be built by the customer. (Source: sources/2026-09-09-aws-validating-multi-region-dr-for-terraform-enterprise-with-aws-fis)

Single-Region by design → customer-operated multi-Region DR

TFE is only supported within a single AWS Region. The multi-Region disaster-recovery pattern documented on this wiki (pilot light, active-passive across us-east-1 / us-west-2, 12–14 min RTO / <1 min RPO) is therefore a customer-designed, customer-operated, customer-tested configuration, not a HashiCorp product feature. Athenahealth built it after an October 2025 us-east-1 event made their single-Region TFE inaccessible; AWS + HashiCorp co-engineered and FIS- validated it. See the source page for the full architecture. (Source: sources/2026-09-09-aws-validating-multi-region-dr-for-terraform-enterprise-with-aws-fis)

TFE-specific DR gotchas

Two internal dependencies deserve explicit handling when designing TFE for multi-Region:

  • TFE encryption password → internal Vault. TFE embeds a HashiCorp Vault; the TFE encryption password protects that Vault's unseal key and root token. A DR instance configured with a different value cannot start or decrypt existing data — so the secret must be replicated to the DR Region (Secrets Manager + KMS) and referenced by the DR launch configuration.
  • Active/Active mode → external Redis. In TFE Active/Active mode, external Redis holds the job queue and cache. A multi-Region design must provision a DR Redis equivalent and explicitly decide what in-flight job loss is acceptable at failover.

TFE exposes a health-check endpoint /_health_check (200 OK when the application is running) that both the NLB target group and the Route 53 health check probe to determine instance/Region health. (Source: sources/2026-09-09-aws-validating-multi-region-dr-for-terraform-enterprise-with-aws-fis)

The state-file circular dependency

TFE's S3-stored state files became the source of the marquee lesson in the AWS post: failover scripts that ran terraform output against the primary-Region state bucket created a circular dependency — during a regional S3 impairment, the failover automation hung on an S3 timeout and could not proceed. The fix is to remove every recovery dependency on the Region being recovered (hardcode identifiers, or use a Region-independent config source). See concepts/circular-dependency for the general shape. (Source: sources/2026-09-09-aws-validating-multi-region-dr-for-terraform-enterprise-with-aws-fis)

Seen in

  • sources/2026-09-09-aws-validating-multi-region-dr-for-terraform-enterprise-with-aws-fis — canonical wiki home. TFE as the single-Region IaC control plane Athenahealth wrapped in a customer-operated active-passive multi-Region DR architecture (Aurora global DB + bidirectional S3 CRR + Route 53 health-check failover), FIS-validated in three phases; source of the encryption-password/Vault + Active/Active-Redis DR gotchas and the state-file circular-dependency lesson.
Last updated · 766 distilled / 2,225 read