SYSTEM Cited by 1 source
AWS Systems Manager Automation¶
AWS Systems Manager (SSM) Automation executes multi-step runbook
documents against AWS resources — each document is an ordered list
of steps (API calls, waits, branches) with onFailure / onCancel
routing. It is AWS's general-purpose operational-automation substrate;
in the resilience-testing context it is the delegation target that
AWS FIS calls to reach resources
FIS cannot natively impair.
Role as an FIS delegation target¶
FIS actions can invoke an SSM Automation document via
ssm:StartAutomationExecution. This is how FIS mutates an SQS queue's
resource policy (a fault with no native FIS action): FIS chains the
document across impairment phases, passing an increasing duration.
(Source: sources/2026-09-09-aws-testing-application-resilience-with-amazon-sqs-and-aws-fault-injection-service)
The SQS-impairment document (4 steps)¶
getTargetQueues— finds SQS queues taggedFIS-Ready: True. CallsListQueuesonce, which returns at most 1,000 queue URLs; in an account with more queues, add pagination or aQueueNamePrefixfilter before relying on it to find every tagged queue.applyDenyAllPolicyToQueues— adds a scopedFISTemporaryDenystatement to each queue's resource policy. Deny only data-plane actions, never management actions, so the automation can remove its own statement during cleanup (see the self-lockout note in systems/aws-iam and concepts/fail-open-vs-fail-closed). Optionally add aDateLessThancondition onaws:CurrentTimeto make the deny self-expiring even if cleanup never runs.waitForDuration— sleeps for the impairment duration (ISO-8601, e.g.PT2M).removeDenyAllPolicyFromQueues— removes theFISTemporaryDenystatement.onFailureandonCancelboth route here so an aborted run still attempts cleanup, and the step raises if it can't restore a policy rather than reporting false success.
The write step is conditioned on aws:ResourceTag/FIS-Ready in the SSM
Automation role's IAM policy, which prevents the automation from touching
untagged queues — a scoping guardrail worth keeping.
(Source: sources/2026-09-09-aws-testing-application-resilience-with-amazon-sqs-and-aws-fault-injection-service)
Concurrency hazard¶
The document reads the queue policy, modifies it, and writes it back — a read-modify-write. Running two impairment experiments against the same queue concurrently can overwrite each other and leave a stale deny behind. Target distinct queues, or run them in sequence. (Source: sources/2026-09-09-aws-testing-application-resilience-with-amazon-sqs-and-aws-fault-injection-service)
Seen in¶
- sources/2026-09-09-aws-testing-application-resilience-with-amazon-sqs-and-aws-fault-injection-service
— canonical wiki home. The 4-step SQS data-plane-deny document FIS
drives; the
onCancel-routed self-cleanup, the tag-conditioned write scope, and the read-modify-write concurrency hazard.
Related¶
- systems/aws-fault-injection-service — the orchestrator that invokes this document.
- systems/aws-iam — the resource-policy evaluation the document manipulates.
- systems/aws-sqs — the target resource.
- concepts/control-plane-data-plane-separation — the document denies data-plane actions, leaves management (control-plane) actions intact.