PATTERN Cited by 1 source
Dead-letter queue¶
Pattern¶
A dead-letter queue (DLQ) is a secondary queue that receives messages a consumer repeatedly fails to process, so they are set aside for investigation instead of blocking the primary queue or being lost. It is a consumer-side safety net for poison messages — messages that will never succeed no matter how many times they are retried (malformed payload, a bug the message reliably triggers, a permanently-missing dependency). Configure a DLQ for every production queue.
Redrive mechanics¶
Redelivery to the DLQ is driven by a maxReceiveCount threshold: a
message moves to the DLQ only after a consumer has received it that
many times without deleting it (i.e. without successfully processing and
acknowledging). Typical maxReceiveCount is 3–5 — enough to ride out
transient failures, few enough to quarantine a genuine poison message
before it wastes much capacity.
Two consequences fall directly out of this mechanic:
- The DLQ depends on the receive path. If
ReceiveMessageitself is blocked (IAM deny, outage, stopped consumers), nothing is delivered, the receive count never increments, and nothing redrives — the DLQ stays empty during the outage and only starts moving on recovery. This is why, in a resilience experiment that denies SQS access, DLQ depth is a recovery-phase signal, not an outage signal. (Source: sources/2026-09-09-aws-testing-application-resilience-with-amazon-sqs-and-aws-fault-injection-service) - Every consumer must tolerate seeing a message more than once.
Redelivery is what feeds a DLQ, so a message that ultimately dead-letters
was processed (or partially processed) up to
maxReceiveCounttimes first. See concepts/at-least-once-delivery and concepts/idempotent-operations.
What a DLQ is not¶
A DLQ is not a producer overflow buffer. A producer that gives up on a send (non-retryable error, or exhausted retry budget) routes nothing to the DLQ — the DLQ only ever receives messages that were enqueued and then failed consumption. Producer-side durability (buffering failed sends to local durable storage and replaying on recovery) is a separate mechanism; conflating the two leaves a gap where dropped sends are lost silently. (Source: sources/2026-09-09-aws-testing-application-resilience-with-amazon-sqs-and-aws-fault-injection-service)
Retention, not depth, is the constraint¶
A DLQ has no depth limit; the binding constraint is the retention period, after which the broker deletes the message (SQS default 4 days, max 14). Practical rules:
- Set the DLQ's retention longer than the source queue's to buy investigation time before messages age out.
- For standard SQS queues the retention clock runs from the original enqueue time and does not reset on the move to the DLQ — time spent in the source queue counts against it. (FIFO queues do reset it.) A message that sat in the source queue for days before dead-lettering may have little DLQ lifetime left.
- After an incident, confirm messages were preserved rather than aged out, and reprocess the DLQ in controlled batches — dumping the whole DLQ back at once can overwhelm a downstream that's still recovering.
Recovery behavior¶
On recovery from an outage, a consumer overwhelmed by the accumulated
backlog can re-fail messages and push some to the DLQ. If healthy
messages (not genuine poison) land there, it's a signal that either
maxReceiveCount is too low or the consumer isn't keeping up with the
backlog — not that the messages are bad. Distinguish genuine poison
messages (reprocess deliberately, in batches, within the retention
window) from backlog-induced re-failures (fix consumer throughput /
raise maxReceiveCount).
(Source: sources/2026-09-09-aws-testing-application-resilience-with-amazon-sqs-and-aws-fault-injection-service)
The DLQ + external re-drive variant¶
A common fallback-tier shape pairs a DLQ with an out-of-band drainer: failed events land in a DLQ, and a scheduled job (cron / Lambda) re-drives them against the downstream until it accepts. Zalando's Order Store uses exactly this — a Lambda publishes to Nakadi, exhausted retries land in the Lambda's built-in SQS DLQ, and a Kubernetes CronJob runs the same publication code to drain it. This re-drive is a second source of duplicates and reordering beyond the initial retry ladder, which is why the consumers must be idempotent (see concepts/at-least-once-delivery).
Related¶
- systems/aws-sqs — the canonical managed DLQ implementation (
maxReceiveCount, redrive policy). - systems/aws-lambda — event-source mappings and async invokes support built-in DLQs.
- concepts/at-least-once-delivery — DLQ redrive is a duplicate source; consumers must be idempotent.
- concepts/idempotent-operations — the consumer-side mandate DLQ redelivery imposes.
- concepts/exponential-backoff-jitter — the in-consumer retry discipline that precedes dead-lettering.
- patterns/circuit-breaker — producer-side complement; the DLQ is the consumer-side safety net.
- patterns/saga-over-long-transaction — where compensating actions, not DLQs, handle multi-step failures.
Seen in¶
- sources/2026-09-09-aws-testing-application-resilience-with-amazon-sqs-and-aws-fault-injection-service
— canonical wiki home. The DLQ mechanics under a resilience
experiment: redrive driven by
maxReceiveCount, DLQ empty during a receive-path outage and only moving on recovery, retention (not depth) as the real constraint, standard-vs-FIFO retention-clock behavior, and "healthy messages in the DLQ =maxReceiveCounttoo low or consumer behind" as a recovery diagnostic.