Skip to content

SYSTEM Cited by 4 sources

AWS Fault Injection Service

AWS Fault Injection Service (AWS FIS) is AWS's managed chaos-engineering service: it orchestrates controlled fault-injection experiments against AWS resources so teams can verify that their recovery mechanisms and observability actually work under failure. It is the managed-cloud analogue of Netflix's Simian Army — where Chaos Monkey kills instances, FIS injects a broader menu of faults (instance termination, API throttling, added latency, AZ power interruption, and — via SSM Automation — arbitrary resource-policy manipulation) as reviewed, stoppable experiment templates.

Core model

  • Experiment template — declares the actions (faults), targets (which resources), durations, and stop conditions. Actions can be chained with startAfter to enforce sequential execution, letting a template escalate severity across phases.
  • Actions — the fault primitives. Some are native (EC2 termination, network latency); others delegate to an SSM Automation document to reach resources FIS doesn't natively impair (e.g. mutating an SQS queue's resource policy).
  • Stop conditions — an FIS experiment halts automatically if a named CloudWatch alarm fires. This is the essential safety control for an escalating experiment. A triggered stop condition also unwinds what it can: FIS cancels the run and any onCancel cleanup step (in a delegated SSM document) executes.
  • Rollback is per-action, not universal. The SQS-deny experiment's onCancel deny-removal makes its rollback clean; but an action like EC2 instance termination has no rollback. Check each action's rollback behavior before relying on a stop condition to limit damage. (Source: sources/2026-09-09-aws-testing-application-resilience-with-amazon-sqs-and-aws-fault-injection-service)

Stop-condition discipline (the load-bearing rule)

Tie the stop condition to a signal that should stay healthy if resilience is working — never to a metric the experiment is designed to move. Alarming on the effect you're injecting (queue depth, send count) trips in the first impairment phase and aborts before longer phases surface anything interesting. Use a customer-impact metric instead: application error rate / failed-transactions, an ALB HTTPCode_Target_5XX_Count, or DLQ depth. Derive the threshold from the hypothesis's recovery window. (Source: sources/2026-09-09-aws-testing-application-resilience-with-amazon-sqs-and-aws-fault-injection-service)

SQS queue-impairment experiment

The canonical worked example: FIS drives an SSM Automation document that applies a scoped deny resource policy to SQS queues tagged FIS-Ready: True, blocking data-plane actions (SendMessage, ReceiveMessage, DeleteMessage, ChangeMessageVisibility, PurgeQueue) while leaving management actions intact — a direct application of concepts/control-plane-data-plane-separation. Four impairment phases (2 / 5 / 7 / 15 min) chained with recovery windows escalate severity to surface fail-fast, backlog, resource-pressure, and systemic failure modes in turn. The experiment tests the application's circuit breakers, backoff, and DLQ behavior — not SQS itself. Blast radius is tuned via the deny's Principal: "*" impairs every caller (service-partition simulation), a role ARN impairs only the app. (Source: sources/2026-09-09-aws-testing-application-resilience-with-amazon-sqs-and-aws-fault-injection-service)

Scenario library

FIS ships scenarios — AWS-authored templates bundling actions, targets, and durations for a recognizable event, so you start from a reviewed definition instead of assembling actions yourself:

  • AZ Availability: Power Interruption — induces the symptoms of losing power in one AZ: zonal EC2/ECS/EKS compute stops, new launches in that AZ fail, subnet connectivity is lost. A sharper test of queue-based decoupling because producers/consumers lose capacity while the queue itself is untargeted — you learn whether surviving consumers absorb the backlog and whether Auto Scaling replaces capacity in the remaining AZs. Defaults to 30 min impairment + 30 min recovery.
  • AZ: Application Slowdown — adds latency between resources within a single AZ, reproducing a partial disruption / grey failure. Useful to validate observability, tune alarm thresholds, and practice AZ-evacuation decisions.

Scenarios carry the same obligations as hand-built templates: write the hypothesis first, and set the stop condition on a customer-impact metric rather than one the scenario is designed to move. (Source: sources/2026-09-09-aws-testing-application-resilience-with-amazon-sqs-and-aws-fault-injection-service)

Progressive DR validation (three phases + combined run)

Beyond single-fault experiments, FIS is used to validate a full multi-Region DR architecture before an event forces it — the discipline of testing failover on your schedule, not the outage's. Athenahealth validated their pilot-light Terraform Enterprise DR in three progressive phases, each building confidence in one layer and each surfacing a bug manual review missed:

  • Phase A — compute. aws:ec2:stop-instances, aws:ec2:asg-insufficient-instance-capacity-error, plus Auto Scaling group suspend/resume. Confirmed 2–3 min ASG replacement — and exposed an outdated AMI reference in the DR launch template, a config drift that only surfaces when the ASG actually launches new instances.
  • Phase B — database. aws:rds:failover-db-cluster on an Aurora global database. Confirmed a 58-second writer promotion — and revealed the application layer didn't meet the hypothesis: TFE connection pooling caused extended reconnection delays (fix: pool timeout 60 s → 10 s). Note this is Aurora's coordinated failover path (requires the primary reachable); an unplanned event uses failover-global-cluster --allow-data-loss.
  • Phase C — S3 connectivity via subnet disruption. See below.
  • Combined — primary-Region impairment. One template suspends the primary ASG, stops instances, holds the faults during a wait window while operators run the failover runbook, then resumes the ASG. Finding no new failure modes beyond the individual phases was itself the confirmation the team wanted. (Source: sources/2026-09-09-aws-validating-multi-region-dr-for-terraform-enterprise-with-aws-fis)

Simulating S3 disruption: there is no direct action

FIS has no direct action to disrupt S3 access. The workaround is aws:network:disrupt-connectivity targeting the compute private subnets: FIS injects network-ACL rules that block egress to S3 service endpoints, simulating regional S3 loss for the instances in those subnets. Important scoping consequence:

  • This tests the S3 consumer (the app losing S3), not S3 cross-Region replication — service-side bucket-to-bucket replication does not traverse your subnet NACLs. To test paused/delayed replication, use the Cross-Region: Connectivity scenario instead.
  • It assumes an in-VPC path to S3 (e.g. a gateway VPC endpoint) so the injected NACL rules actually sit on the egress path.

In the TFE case this phase confirmed <30 s replication lag and exposed the state-file circular dependency in the failover scripts — the highest-value finding of the engagement. (Source: sources/2026-09-09-aws-validating-multi-region-dr-for-terraform-enterprise-with-aws-fis)

Relationship to Resilience Hub

FIS is the execution layer; systems/aws-resilience-hub is the orchestration layer. Resilience Hub stores RTO/RPO targets that the test-generation layer uses to design FIS experiments, and correlates FIS experiment outcomes with resilience policies during gap analysis. (Source: sources/2026-06-22-aws-architecting-ai-powered-resilience-framework-on-aws)

Seen in

Last updated · 766 distilled / 2,225 read