Skip to content

PATTERN Cited by 1 source

Dead-letter queue

Pattern

A dead-letter queue (DLQ) is a secondary queue that receives messages a consumer repeatedly fails to process, so they are set aside for investigation instead of blocking the primary queue or being lost. It is a consumer-side safety net for poison messages — messages that will never succeed no matter how many times they are retried (malformed payload, a bug the message reliably triggers, a permanently-missing dependency). Configure a DLQ for every production queue.

Redrive mechanics

Redelivery to the DLQ is driven by a maxReceiveCount threshold: a message moves to the DLQ only after a consumer has received it that many times without deleting it (i.e. without successfully processing and acknowledging). Typical maxReceiveCount is 3–5 — enough to ride out transient failures, few enough to quarantine a genuine poison message before it wastes much capacity.

Two consequences fall directly out of this mechanic:

What a DLQ is not

A DLQ is not a producer overflow buffer. A producer that gives up on a send (non-retryable error, or exhausted retry budget) routes nothing to the DLQ — the DLQ only ever receives messages that were enqueued and then failed consumption. Producer-side durability (buffering failed sends to local durable storage and replaying on recovery) is a separate mechanism; conflating the two leaves a gap where dropped sends are lost silently. (Source: sources/2026-09-09-aws-testing-application-resilience-with-amazon-sqs-and-aws-fault-injection-service)

Retention, not depth, is the constraint

A DLQ has no depth limit; the binding constraint is the retention period, after which the broker deletes the message (SQS default 4 days, max 14). Practical rules:

  • Set the DLQ's retention longer than the source queue's to buy investigation time before messages age out.
  • For standard SQS queues the retention clock runs from the original enqueue time and does not reset on the move to the DLQ — time spent in the source queue counts against it. (FIFO queues do reset it.) A message that sat in the source queue for days before dead-lettering may have little DLQ lifetime left.
  • After an incident, confirm messages were preserved rather than aged out, and reprocess the DLQ in controlled batches — dumping the whole DLQ back at once can overwhelm a downstream that's still recovering.

Recovery behavior

On recovery from an outage, a consumer overwhelmed by the accumulated backlog can re-fail messages and push some to the DLQ. If healthy messages (not genuine poison) land there, it's a signal that either maxReceiveCount is too low or the consumer isn't keeping up with the backlog — not that the messages are bad. Distinguish genuine poison messages (reprocess deliberately, in batches, within the retention window) from backlog-induced re-failures (fix consumer throughput / raise maxReceiveCount). (Source: sources/2026-09-09-aws-testing-application-resilience-with-amazon-sqs-and-aws-fault-injection-service)

The DLQ + external re-drive variant

A common fallback-tier shape pairs a DLQ with an out-of-band drainer: failed events land in a DLQ, and a scheduled job (cron / Lambda) re-drives them against the downstream until it accepts. Zalando's Order Store uses exactly this — a Lambda publishes to Nakadi, exhausted retries land in the Lambda's built-in SQS DLQ, and a Kubernetes CronJob runs the same publication code to drain it. This re-drive is a second source of duplicates and reordering beyond the initial retry ladder, which is why the consumers must be idempotent (see concepts/at-least-once-delivery).

Seen in

Last updated · 766 distilled / 2,225 read