Testing application resilience with Amazon SQS and AWS Fault Injection Service¶
Summary¶
An AWS Architecture Blog post on running a resilience experiment
against an SQS-decoupled producer/consumer application using
AWS FIS and
SSM Automation. The
experiment injects a scoped deny resource policy on the queue that
blocks only data-plane operations (SendMessage, ReceiveMessage,
DeleteMessage, ChangeMessageVisibility, PurgeQueue) while leaving
queue-management actions untouched, then removes it and observes
recovery. The core reframing: the experiment tests your
application's recovery mechanisms and your observability, not SQS
itself. The deny simulates what the app would experience during a
network partition, IAM change, or bad deploy. Impairment escalates
across four phases (2 / 5 / 7 / 15 minutes) with recovery windows
between, each duration designed to surface a different class of failure
— fail-fast + circuit-breaker activation, backlog accumulation,
thread-pool/memory pressure, and systemic limits under prolonged loss.
Key takeaways¶
-
The goal is not to verify SQS works — it's to learn what your app does when SQS fails, and whether you'd notice. A resilience experiment tests recovery mechanisms and observability at once; every gap between the written hypothesis and the observed result is the work list. (Source: sources/2026-09-09-aws-testing-application-resilience-with-amazon-sqs-and-aws-fault-injection-service)
-
Never deny
sqs:*— an explicit IAM Deny overrides every Allow, including the automation's own permission to remove the policy later. A deny coveringsqs:SetQueueAttributes/sqs:AddPermission/sqs:RemovePermissioncan lock the queue so even the role that applied it can't clean up. Scope the deny to data-plane actions only. This is a concrete instance of IAM's explicit-deny-wins evaluation rule creating a self-inflicted lockout — see concepts/fail-open-vs-fail-closed and systems/aws-iam. Defensive add-ons: a validation step in the automation that refuses any deny covering management actions, and a self-expiringDateLessThancondition onaws:CurrentTimeso the deny stops applying even if cleanup never runs. -
Scope the blast radius with
Principal."Principal": "*"denies the data plane to every caller (app, admins, canaries) — faithfully simulating a service partition but impairing more than your app on a shared queue. Scoping to the application's IAM role ("Principal": {"AWS": "arn:aws:iam::<acct>:role/<app-role>"}) matches all sessions of that role and impairs only the app; on a shared queue this blends impaired and healthy traffic in queue-level metrics, so you read recovery as a return to pre-event levels rather than to zero-and-back. (See concepts/blast-radius.) -
Producer side and consumer side fail differently — watch them separately. Producer (
SendMessage): healthy = recognizesAccessDenied, fails fast, opens circuit breaker after 3–5 consecutive failures, buffers to durable fallback storage rather than dropping. Signal:NumberOfMessagesSent→ ~0, circuit-breaker state metric, producer CPU/threads. Consumer (ReceiveMessage/DeleteMessage): healthy = backs off its poll loop, in-flight messages return after visibility timeout. Signal:ApproximateNumberOfMessagesVisiblegoes flat/climbs,ApproximateAgeOfOldestMessageclimbs. -
A 403 (
AccessDenied) is non-retryable — this experiment validates that your code recognizes it and stops, NOT your backoff/retry path. To exercise retries and backoff, inject a retryable fault (throttling, timeouts) instead. Retry only errors that might succeed on repeat; fail fast on non-retryable ones. Configure the AWS SDK's built-in retry behavior rather than rolling your own. -
Don't expect the DLQ to fill during impairment — redrive is driven by
maxReceiveCount, which requires a consumer to receive a message that many times. WithReceiveMessagedenied, nothing is delivered, the receive count doesn't increment, and nothing redrives. The DLQ depends on the very call that's blocked, so it's a recovery-phase signal, not an outage signal. On recovery, a consumer overwhelmed by the accumulated backlog can re-fail messages and push healthy ones to the DLQ — if healthy messages land there,maxReceiveCountis too low or the consumer isn't keeping up. See patterns/dead-letter-queue. -
The stop condition must alarm on a signal that should stay healthy if resilience is working — never on a metric the experiment is designed to move. Alarming on
ApproximateAgeOfOldestMessageorNumberOfMessagesSentwould trip in the first 2-minute phase and abort before the longer phases surface anything — "you'd be alarming on the effect you're injecting." Instead tie the stop condition to a customer-impact metric (application error rate / failed-orders-per- minute, ALBHTTPCode_Target_5XX_Count, or DLQ depth). Derive the threshold from the hypothesis's recovery window (e.g. "failed orders10 for 2 consecutive 1-minute periods"). A triggered stop condition unwinds via the automation's
onCancelstep (removes the deny) — but that rollback is a property of this experiment's design, not of stop conditions in general (EC2 termination has no rollback). -
Write the hypothesis before you run. Two shapes: a prediction hypothesis when you already have resilience patterns ("we expect the producer to open its circuit breaker within 30s, fail fast, and buffer to durable storage; on recovery replay the buffer and return to baseline within 2 min") or a discovery hypothesis for a never-tested failure mode ("we'll block access for 2 min and observe how the producer handles failed sends and whether the consumer recovers unaided"). Either way, state the metrics you'll judge it by.
-
Preservation and relevance are different questions. After a long outage some buffered/backlogged messages represent requests the client has already given up on — processing them spends recovery capacity on stale intent. Compare each message's timestamp to now as you consume, and drop or sideline stale work deliberately rather than letting it age out. (AWS Well-Architected REL05-BP04.)
Systems / concepts / patterns extracted¶
- Systems: systems/aws-fault-injection-service (the orchestrator), systems/aws-sqs (target queue + DLQ), systems/aws-systems-manager-automation (applies/removes the deny policy via a 4-step document), systems/aws-iam (scoped deny resource policy; explicit-deny-wins), systems/aws-cloudwatch (queue + application metrics, dashboard, stop-condition alarm), systems/aws-resilience-hub (posture assessment / next step), systems/resilience4j (circuit-breaker implementation implied by the producer behaviors), systems/aws-application-load-balancer (5xx proxy metric for stop condition), systems/aws-sdk-retry (built-in configurable retry).
- Concepts: concepts/chaos-engineering, concepts/control-plane-data-plane-separation (the deny targets data-plane actions, leaves control-plane/management intact), concepts/at-least-once-delivery (redelivery + idempotency mandate), concepts/idempotent-operations, concepts/exponential-backoff-jitter, concepts/backpressure (poll-loop backoff, bounded in-flight), concepts/graceful-degradation, concepts/blast-radius (Principal scoping), concepts/grey-failure (the AZ-application-slowdown scenario), concepts/fail-open-vs-fail-closed (deny-lockout).
- Patterns: patterns/circuit-breaker, patterns/dead-letter-queue, patterns/staged-rollout (the progressive 2→5→7→15-minute escalation is a staged-severity rollout of the fault itself).
Operational numbers / details¶
- Four-phase escalation: Impair 2 min → Recover 3 min → Impair 5 min
→ Recover 3 min → Impair 7 min → Recover 2 min → Impair 15 min.
startAfterfields enforce sequential execution. - What each duration surfaces: 2 min = fail-fast + circuit-breaker activation; 5 min = backlog accumulation; 7 min = thread-pool / memory pressure; 15 min = systemic limits under prolonged loss.
- SSM Automation document (4 steps):
getTargetQueues(finds queues taggedFIS-Ready: TrueviaListQueues, capped at 1,000 URLs — add pagination or aQueueNamePrefixfilter beyond that) →applyDenyAllPolicyToQueues(adds scopedFISTemporaryDenystatement) →waitForDuration(ISO-8601, e.g.PT2M) →removeDenyAllPolicyFromQueues.onFailure/onCancelroute to the removal step, which raises if it can't restore rather than falsely reporting success. - IAM roles: FIS role trusted by
fis.amazonaws.comwithssm:StartAutomationExecution+iam:PassRole; SSM Automation role withsqs:GetQueueAttributes/SetQueueAttributes/ListQueues/ListQueueTags, conditioned onaws:ResourceTag/FIS-Ready. - DLQ mechanics: no depth limit; constraint is the retention period
(default 4 days, max 14). Standard queues run the retention clock from
original enqueue time — time in the source queue counts against it;
FIFO queues reset it. Set DLQ retention longer than the source queue's.
maxReceiveCounttypically 3–5. - Circuit-breaker recovery: half-opens then closes within ~30 s on a clean run; a clean square-wave on the state metric that lags each fault window slightly.
- Code: SSM Automation doc, FIS experiment template, and example IAM
policies published in the
aws-samples/fis-template-library(sqs-queue-impairment) on GitHub. - Expansion paths: partial failure (impair a subset of queues — don't
run two impairments against the same queue concurrently, the
read-modify-write races and can leave a stale deny); consumer-side-only
(deny only
ReceiveMessage+DeleteMessage); compound failures (run alongside EC2 termination / latency injection); FIS scenarios (AZ Availability: Power Interruption; AZ: Application Slowdown — a grey failure).
Caveats¶
- The example workload behaviors (circuit breakers, local buffering, thread pools) assume long-running producer/consumer services, not short-lived Lambda invocations.
- The application under test is the one prerequisite with no example in the repo — the library ships the fault injection, not the workload. Without app instrumentation (circuit-breaker state, dropped-message counters, fallback-store writes, duplicate-processing metrics) you only watch queue metrics move and learn little.
- Count-based metrics (
NumberOfMessagesSent/Received/Deleted) include retries and duplicates — treat them as trend indicators, not exact unique-message counts. - Run in non-production first; in production, confirm change-management approvals and have a documented rollback plan.
Source¶
- Original: https://aws.amazon.com/blogs/architecture/testing-application-resilience-with-amazon-sqs-and-aws-fault-injection-service/
- Raw markdown:
raw/aws/2026-09-09-testing-application-resilience-with-amazon-sqs-and-aws-fault-741f3d22.md
Related¶
- systems/aws-fault-injection-service — the fault-injection orchestrator this post exercises.
- systems/aws-sqs — the decoupling primitive under test.
- patterns/circuit-breaker — the producer-side mechanism the 2-minute phase validates.
- patterns/dead-letter-queue — the consumer-side safety net; a recovery-phase signal.
- concepts/chaos-engineering — the discipline this experiment belongs to.
- concepts/control-plane-data-plane-separation — the deny targets data-plane actions only.
- concepts/at-least-once-delivery — redelivery + idempotency mandate on recovery.