SYSTEM Cited by 4 sources
AWS Fault Injection Service¶
AWS Fault Injection Service (AWS FIS) is AWS's managed chaos-engineering service: it orchestrates controlled fault-injection experiments against AWS resources so teams can verify that their recovery mechanisms and observability actually work under failure. It is the managed-cloud analogue of Netflix's Simian Army — where Chaos Monkey kills instances, FIS injects a broader menu of faults (instance termination, API throttling, added latency, AZ power interruption, and — via SSM Automation — arbitrary resource-policy manipulation) as reviewed, stoppable experiment templates.
Core model¶
- Experiment template — declares the actions (faults), targets
(which resources), durations, and stop conditions. Actions can be
chained with
startAfterto enforce sequential execution, letting a template escalate severity across phases. - Actions — the fault primitives. Some are native (EC2 termination, network latency); others delegate to an SSM Automation document to reach resources FIS doesn't natively impair (e.g. mutating an SQS queue's resource policy).
- Stop conditions — an FIS experiment halts automatically if a
named CloudWatch alarm fires. This is the essential safety control for
an escalating experiment. A triggered stop condition also unwinds
what it can: FIS cancels the run and any
onCancelcleanup step (in a delegated SSM document) executes. - Rollback is per-action, not universal. The SQS-deny experiment's
onCanceldeny-removal makes its rollback clean; but an action like EC2 instance termination has no rollback. Check each action's rollback behavior before relying on a stop condition to limit damage. (Source: sources/2026-09-09-aws-testing-application-resilience-with-amazon-sqs-and-aws-fault-injection-service)
Stop-condition discipline (the load-bearing rule)¶
Tie the stop condition to a signal that should stay healthy if
resilience is working — never to a metric the experiment is designed to
move. Alarming on the effect you're injecting (queue depth, send
count) trips in the first impairment phase and aborts before longer
phases surface anything interesting. Use a customer-impact metric
instead: application error rate / failed-transactions, an ALB
HTTPCode_Target_5XX_Count, or DLQ depth. Derive the threshold from the
hypothesis's recovery window.
(Source: sources/2026-09-09-aws-testing-application-resilience-with-amazon-sqs-and-aws-fault-injection-service)
SQS queue-impairment experiment¶
The canonical worked example: FIS drives an SSM Automation document that
applies a scoped deny resource policy to SQS queues tagged
FIS-Ready: True, blocking data-plane actions (SendMessage,
ReceiveMessage, DeleteMessage, ChangeMessageVisibility,
PurgeQueue) while leaving management actions intact — a direct
application of concepts/control-plane-data-plane-separation. Four
impairment phases (2 / 5 / 7 / 15 min) chained with recovery windows
escalate severity to surface fail-fast, backlog, resource-pressure, and
systemic failure modes in turn. The experiment tests the application's
circuit breakers,
backoff, and
DLQ behavior — not SQS itself.
Blast radius is tuned via the deny's Principal: "*" impairs every
caller (service-partition simulation), a role ARN impairs only the app.
(Source: sources/2026-09-09-aws-testing-application-resilience-with-amazon-sqs-and-aws-fault-injection-service)
Scenario library¶
FIS ships scenarios — AWS-authored templates bundling actions, targets, and durations for a recognizable event, so you start from a reviewed definition instead of assembling actions yourself:
- AZ Availability: Power Interruption — induces the symptoms of losing power in one AZ: zonal EC2/ECS/EKS compute stops, new launches in that AZ fail, subnet connectivity is lost. A sharper test of queue-based decoupling because producers/consumers lose capacity while the queue itself is untargeted — you learn whether surviving consumers absorb the backlog and whether Auto Scaling replaces capacity in the remaining AZs. Defaults to 30 min impairment + 30 min recovery.
- AZ: Application Slowdown — adds latency between resources within a single AZ, reproducing a partial disruption / grey failure. Useful to validate observability, tune alarm thresholds, and practice AZ-evacuation decisions.
Scenarios carry the same obligations as hand-built templates: write the hypothesis first, and set the stop condition on a customer-impact metric rather than one the scenario is designed to move. (Source: sources/2026-09-09-aws-testing-application-resilience-with-amazon-sqs-and-aws-fault-injection-service)
Progressive DR validation (three phases + combined run)¶
Beyond single-fault experiments, FIS is used to validate a full multi-Region DR architecture before an event forces it — the discipline of testing failover on your schedule, not the outage's. Athenahealth validated their pilot-light Terraform Enterprise DR in three progressive phases, each building confidence in one layer and each surfacing a bug manual review missed:
- Phase A — compute.
aws:ec2:stop-instances,aws:ec2:asg-insufficient-instance-capacity-error, plus Auto Scaling group suspend/resume. Confirmed 2–3 min ASG replacement — and exposed an outdated AMI reference in the DR launch template, a config drift that only surfaces when the ASG actually launches new instances. - Phase B — database.
aws:rds:failover-db-clusteron an Aurora global database. Confirmed a 58-second writer promotion — and revealed the application layer didn't meet the hypothesis: TFE connection pooling caused extended reconnection delays (fix: pool timeout 60 s → 10 s). Note this is Aurora's coordinated failover path (requires the primary reachable); an unplanned event usesfailover-global-cluster --allow-data-loss. - Phase C — S3 connectivity via subnet disruption. See below.
- Combined — primary-Region impairment. One template suspends the primary ASG, stops instances, holds the faults during a wait window while operators run the failover runbook, then resumes the ASG. Finding no new failure modes beyond the individual phases was itself the confirmation the team wanted. (Source: sources/2026-09-09-aws-validating-multi-region-dr-for-terraform-enterprise-with-aws-fis)
Simulating S3 disruption: there is no direct action¶
FIS has no direct action to disrupt S3 access. The workaround is
aws:network:disrupt-connectivity targeting the compute private
subnets: FIS injects network-ACL rules that block egress to S3 service
endpoints, simulating regional S3 loss for the instances in those
subnets. Important scoping consequence:
- This tests the S3 consumer (the app losing S3), not S3 cross-Region replication — service-side bucket-to-bucket replication does not traverse your subnet NACLs. To test paused/delayed replication, use the Cross-Region: Connectivity scenario instead.
- It assumes an in-VPC path to S3 (e.g. a gateway VPC endpoint) so the injected NACL rules actually sit on the egress path.
In the TFE case this phase confirmed <30 s replication lag and exposed the state-file circular dependency in the failover scripts — the highest-value finding of the engagement. (Source: sources/2026-09-09-aws-validating-multi-region-dr-for-terraform-enterprise-with-aws-fis)
Relationship to Resilience Hub¶
FIS is the execution layer; systems/aws-resilience-hub is the orchestration layer. Resilience Hub stores RTO/RPO targets that the test-generation layer uses to design FIS experiments, and correlates FIS experiment outcomes with resilience policies during gap analysis. (Source: sources/2026-06-22-aws-architecting-ai-powered-resilience-framework-on-aws)
Seen in¶
- sources/2026-09-30-aws-running-multi-day-az-evacuation-drills-with-arc-zonal-shift — FIS as the injection sibling of ARC Zonal Shift. The AZ-evacuation drill is a planned traffic shift (ARC is the response); FIS injects the impairment that validates the response — its AZ Availability: Power Interruption scenario reproduces an AZ loss, and FIS has a dedicated recovery action for zonal autoshift so the automatic-shift path can itself be chaos-tested. Same discipline as the manual drill at the next rung of automation (AWS Architecture Blog, 2026-09-30).
- sources/2026-09-09-aws-validating-multi-region-dr-for-terraform-enterprise-with-aws-fis
— DR-validation use case. Three progressive FIS phases (EC2/ASG,
aws:rds:failover-db-clusterAurora failover,aws:network:disrupt-connectivityS3-consumer disruption) plus a combined primary-Region-impairment run, validating a customer-operated multi-Region TFE DR. Canonical example of FIS finding config drift (stale DR AMI), an app-layer miss (TFE connection-pool timeout), and a circular dependency that only surfaces under real disruption. Also the canonical "no direct S3 action → disrupt the subnet's NACLs" workaround. - sources/2026-09-09-aws-testing-application-resilience-with-amazon-sqs-and-aws-fault-injection-service
— canonical wiki home. FIS + SSM Automation running a four-phase
SQS data-plane-deny experiment; the stop-condition discipline
(alarm on customer impact, not on the injected effect), the
Principal-scoped blast radius, and the scenario library (AZ Power Interruption, AZ Application Slowdown). - sources/2026-06-22-aws-architecting-ai-powered-resilience-framework-on-aws — FIS as the fault-injection execution layer beneath Resilience Hub in AWS's five-layer AI-powered resilience framework.
Related¶
- systems/aws-resilience-hub — the orchestration/assessment layer above FIS.
- systems/aws-systems-manager-automation — the delegation substrate FIS uses to reach non-native resources.
- systems/netflix-simian-army, systems/netflix-chaos-monkey — the pre-cloud-managed lineage of the same discipline.
- concepts/chaos-engineering — the discipline FIS operationalizes.
- concepts/blast-radius — tuned via the deny
Principal. - patterns/staged-rollout — the escalating-severity phase structure applied to the fault.