Skip to content

SYSTEM Cited by 1 source

AWS Systems Manager Automation

AWS Systems Manager (SSM) Automation executes multi-step runbook documents against AWS resources — each document is an ordered list of steps (API calls, waits, branches) with onFailure / onCancel routing. It is AWS's general-purpose operational-automation substrate; in the resilience-testing context it is the delegation target that AWS FIS calls to reach resources FIS cannot natively impair.

Role as an FIS delegation target

FIS actions can invoke an SSM Automation document via ssm:StartAutomationExecution. This is how FIS mutates an SQS queue's resource policy (a fault with no native FIS action): FIS chains the document across impairment phases, passing an increasing duration. (Source: sources/2026-09-09-aws-testing-application-resilience-with-amazon-sqs-and-aws-fault-injection-service)

The SQS-impairment document (4 steps)

  1. getTargetQueues — finds SQS queues tagged FIS-Ready: True. Calls ListQueues once, which returns at most 1,000 queue URLs; in an account with more queues, add pagination or a QueueNamePrefix filter before relying on it to find every tagged queue.
  2. applyDenyAllPolicyToQueues — adds a scoped FISTemporaryDeny statement to each queue's resource policy. Deny only data-plane actions, never management actions, so the automation can remove its own statement during cleanup (see the self-lockout note in systems/aws-iam and concepts/fail-open-vs-fail-closed). Optionally add a DateLessThan condition on aws:CurrentTime to make the deny self-expiring even if cleanup never runs.
  3. waitForDuration — sleeps for the impairment duration (ISO-8601, e.g. PT2M).
  4. removeDenyAllPolicyFromQueues — removes the FISTemporaryDeny statement. onFailure and onCancel both route here so an aborted run still attempts cleanup, and the step raises if it can't restore a policy rather than reporting false success.

The write step is conditioned on aws:ResourceTag/FIS-Ready in the SSM Automation role's IAM policy, which prevents the automation from touching untagged queues — a scoping guardrail worth keeping. (Source: sources/2026-09-09-aws-testing-application-resilience-with-amazon-sqs-and-aws-fault-injection-service)

Concurrency hazard

The document reads the queue policy, modifies it, and writes it back — a read-modify-write. Running two impairment experiments against the same queue concurrently can overwrite each other and leave a stale deny behind. Target distinct queues, or run them in sequence. (Source: sources/2026-09-09-aws-testing-application-resilience-with-amazon-sqs-and-aws-fault-injection-service)

Seen in

Last updated · 766 distilled / 2,225 read