Skip to content

AWS 2026-09-09

Read original ↗

Testing application resilience with Amazon SQS and AWS Fault Injection Service

Summary

An AWS Architecture Blog post on running a resilience experiment against an SQS-decoupled producer/consumer application using AWS FIS and SSM Automation. The experiment injects a scoped deny resource policy on the queue that blocks only data-plane operations (SendMessage, ReceiveMessage, DeleteMessage, ChangeMessageVisibility, PurgeQueue) while leaving queue-management actions untouched, then removes it and observes recovery. The core reframing: the experiment tests your application's recovery mechanisms and your observability, not SQS itself. The deny simulates what the app would experience during a network partition, IAM change, or bad deploy. Impairment escalates across four phases (2 / 5 / 7 / 15 minutes) with recovery windows between, each duration designed to surface a different class of failure — fail-fast + circuit-breaker activation, backlog accumulation, thread-pool/memory pressure, and systemic limits under prolonged loss.

Key takeaways

  1. The goal is not to verify SQS works — it's to learn what your app does when SQS fails, and whether you'd notice. A resilience experiment tests recovery mechanisms and observability at once; every gap between the written hypothesis and the observed result is the work list. (Source: sources/2026-09-09-aws-testing-application-resilience-with-amazon-sqs-and-aws-fault-injection-service)

  2. Never deny sqs:* — an explicit IAM Deny overrides every Allow, including the automation's own permission to remove the policy later. A deny covering sqs:SetQueueAttributes / sqs:AddPermission / sqs:RemovePermission can lock the queue so even the role that applied it can't clean up. Scope the deny to data-plane actions only. This is a concrete instance of IAM's explicit-deny-wins evaluation rule creating a self-inflicted lockout — see concepts/fail-open-vs-fail-closed and systems/aws-iam. Defensive add-ons: a validation step in the automation that refuses any deny covering management actions, and a self-expiring DateLessThan condition on aws:CurrentTime so the deny stops applying even if cleanup never runs.

  3. Scope the blast radius with Principal. "Principal": "*" denies the data plane to every caller (app, admins, canaries) — faithfully simulating a service partition but impairing more than your app on a shared queue. Scoping to the application's IAM role ("Principal": {"AWS": "arn:aws:iam::<acct>:role/<app-role>"}) matches all sessions of that role and impairs only the app; on a shared queue this blends impaired and healthy traffic in queue-level metrics, so you read recovery as a return to pre-event levels rather than to zero-and-back. (See concepts/blast-radius.)

  4. Producer side and consumer side fail differently — watch them separately. Producer (SendMessage): healthy = recognizes AccessDenied, fails fast, opens circuit breaker after 3–5 consecutive failures, buffers to durable fallback storage rather than dropping. Signal: NumberOfMessagesSent → ~0, circuit-breaker state metric, producer CPU/threads. Consumer (ReceiveMessage / DeleteMessage): healthy = backs off its poll loop, in-flight messages return after visibility timeout. Signal: ApproximateNumberOfMessagesVisible goes flat/climbs, ApproximateAgeOfOldestMessage climbs.

  5. A 403 (AccessDenied) is non-retryable — this experiment validates that your code recognizes it and stops, NOT your backoff/retry path. To exercise retries and backoff, inject a retryable fault (throttling, timeouts) instead. Retry only errors that might succeed on repeat; fail fast on non-retryable ones. Configure the AWS SDK's built-in retry behavior rather than rolling your own.

  6. Don't expect the DLQ to fill during impairment — redrive is driven by maxReceiveCount, which requires a consumer to receive a message that many times. With ReceiveMessage denied, nothing is delivered, the receive count doesn't increment, and nothing redrives. The DLQ depends on the very call that's blocked, so it's a recovery-phase signal, not an outage signal. On recovery, a consumer overwhelmed by the accumulated backlog can re-fail messages and push healthy ones to the DLQ — if healthy messages land there, maxReceiveCount is too low or the consumer isn't keeping up. See patterns/dead-letter-queue.

  7. The stop condition must alarm on a signal that should stay healthy if resilience is working — never on a metric the experiment is designed to move. Alarming on ApproximateAgeOfOldestMessage or NumberOfMessagesSent would trip in the first 2-minute phase and abort before the longer phases surface anything — "you'd be alarming on the effect you're injecting." Instead tie the stop condition to a customer-impact metric (application error rate / failed-orders-per- minute, ALB HTTPCode_Target_5XX_Count, or DLQ depth). Derive the threshold from the hypothesis's recovery window (e.g. "failed orders

    10 for 2 consecutive 1-minute periods"). A triggered stop condition unwinds via the automation's onCancel step (removes the deny) — but that rollback is a property of this experiment's design, not of stop conditions in general (EC2 termination has no rollback).

  8. Write the hypothesis before you run. Two shapes: a prediction hypothesis when you already have resilience patterns ("we expect the producer to open its circuit breaker within 30s, fail fast, and buffer to durable storage; on recovery replay the buffer and return to baseline within 2 min") or a discovery hypothesis for a never-tested failure mode ("we'll block access for 2 min and observe how the producer handles failed sends and whether the consumer recovers unaided"). Either way, state the metrics you'll judge it by.

  9. Preservation and relevance are different questions. After a long outage some buffered/backlogged messages represent requests the client has already given up on — processing them spends recovery capacity on stale intent. Compare each message's timestamp to now as you consume, and drop or sideline stale work deliberately rather than letting it age out. (AWS Well-Architected REL05-BP04.)

Systems / concepts / patterns extracted

Operational numbers / details

  • Four-phase escalation: Impair 2 min → Recover 3 min → Impair 5 min → Recover 3 min → Impair 7 min → Recover 2 min → Impair 15 min. startAfter fields enforce sequential execution.
  • What each duration surfaces: 2 min = fail-fast + circuit-breaker activation; 5 min = backlog accumulation; 7 min = thread-pool / memory pressure; 15 min = systemic limits under prolonged loss.
  • SSM Automation document (4 steps): getTargetQueues (finds queues tagged FIS-Ready: True via ListQueues, capped at 1,000 URLs — add pagination or a QueueNamePrefix filter beyond that) → applyDenyAllPolicyToQueues (adds scoped FISTemporaryDeny statement) → waitForDuration (ISO-8601, e.g. PT2M) → removeDenyAllPolicyFromQueues. onFailure/onCancel route to the removal step, which raises if it can't restore rather than falsely reporting success.
  • IAM roles: FIS role trusted by fis.amazonaws.com with ssm:StartAutomationExecution + iam:PassRole; SSM Automation role with sqs:GetQueueAttributes / SetQueueAttributes / ListQueues / ListQueueTags, conditioned on aws:ResourceTag/FIS-Ready.
  • DLQ mechanics: no depth limit; constraint is the retention period (default 4 days, max 14). Standard queues run the retention clock from original enqueue time — time in the source queue counts against it; FIFO queues reset it. Set DLQ retention longer than the source queue's. maxReceiveCount typically 3–5.
  • Circuit-breaker recovery: half-opens then closes within ~30 s on a clean run; a clean square-wave on the state metric that lags each fault window slightly.
  • Code: SSM Automation doc, FIS experiment template, and example IAM policies published in the aws-samples/fis-template-library (sqs-queue-impairment) on GitHub.
  • Expansion paths: partial failure (impair a subset of queues — don't run two impairments against the same queue concurrently, the read-modify-write races and can leave a stale deny); consumer-side-only (deny only ReceiveMessage + DeleteMessage); compound failures (run alongside EC2 termination / latency injection); FIS scenarios (AZ Availability: Power Interruption; AZ: Application Slowdown — a grey failure).

Caveats

  • The example workload behaviors (circuit breakers, local buffering, thread pools) assume long-running producer/consumer services, not short-lived Lambda invocations.
  • The application under test is the one prerequisite with no example in the repo — the library ships the fault injection, not the workload. Without app instrumentation (circuit-breaker state, dropped-message counters, fallback-store writes, duplicate-processing metrics) you only watch queue metrics move and learn little.
  • Count-based metrics (NumberOfMessagesSent/Received/Deleted) include retries and duplicates — treat them as trend indicators, not exact unique-message counts.
  • Run in non-production first; in production, confirm change-management approvals and have a documented rollback plan.

Source

Last updated · 766 distilled / 2,225 read