CONCEPT Cited by 3 sources
At-least-once delivery¶
Definition¶
At-least-once delivery is the messaging guarantee that each message will eventually be delivered to a consumer at least once, with no upper bound on how many duplicates the consumer may observe. The system prefers repeating a message to losing one. Most managed message buses — Kafka, SQS, DynamoDB Streams → Lambda, Nakadi, SNS fan-out — pick this as their default guarantee because it's achievable with simple broker logic and survives broker crashes, network partitions, and consumer failures.
The three classic guarantees¶
| Guarantee | Delivery | Loss risk | Duplicate risk |
|---|---|---|---|
| At-most-once | ≤ 1 copy | yes (on crash) | no |
| At-least-once | ≥ 1 copy | no | yes |
| Exactly-once | exactly 1 copy | no | no |
Exactly-once is expensive — it requires transactional producers and consumers, coordinated commits, and deduplication state. Most production systems choose at-least-once and push idempotency to the consumer.
Where duplicates come from¶
Even a well-behaved broker will produce duplicates at the consumer in any of these scenarios:
- Consumer crashes after processing but before acknowledging. Broker redelivers on restart.
- Acknowledgement is lost in transit. Network blip between consumer and broker.
- Producer retries after timeout. Producer got no response, re-sent the message; broker accepted both copies.
- DLQ requeue. A dead-letter handler drains the DLQ and re-publishes; consumers see the event twice (once for the original retry chain, once from the requeue).
- Partition rebalance mid-flight. The new owner starts from the last committed offset, which may be behind the last processed.
The consumer-side mandate: idempotency¶
Under at-least-once, every consumer must tolerate seeing a message more than once without side effects. Common idempotency patterns:
- Dedup by message ID — track seen IDs in a bounded-time key- value store; drop duplicates.
- Idempotent writes — model the downstream write as "set state to X" not "increment by 1"; replays converge.
- Upserts with logical timestamps — last-write-wins by client- supplied monotonic clock.
- Write + compare — compute the effect; compare to current state; only apply the delta.
Which one fits depends on what the consumer does. Event-sourcing consumers naturally do the last one; cache-update consumers fit the idempotent-write form; notification senders use dedup-by-ID.
The Zalando framing¶
Zalando's Lambda relay publishing events to Nakadi is explicitly at-least-once: on transient Nakadi failure, the Lambda retries with exponential backoff; on exhaustion, the event goes to an SQS DLQ; a Kubernetes CronJob requeues from DLQ until Nakadi accepts. The post acknowledges both the ordering and duplication consequences:
"In case the publication to Nakadi fails, e.g. due to timeouts, the request is retried. If all the retries fail then we make use of an AWS SQS queue as fallback storage... This also means that we do not guarantee that the events are published in the correct order." (Source: )
Consumers of those Nakadi events must handle both replays and reordering.
Ordering vs delivery¶
At-least-once and per-key ordering are orthogonal:
- Some systems preserve per-key order on redelivery (Kafka by partition, DynamoDB Streams by item hash key).
- Some don't (any system with DLQ requeue or parallel retry workers — as in Zalando's design).
If the consumer needs stable order, it has to either: (a) pick a substrate that preserves order and accept serial processing per key; or (b) reconstruct order via embedded sequence numbers / logical timestamps in the payload.
Seen in¶
- — Zalando's transactional-outbox relay: DynamoDB Streams → Lambda → Nakadi is at-least-once by construction (Lambda retries
-
DLQ + CronJob requeue); the post explicitly flags out-of-order delivery as the accepted trade-off. Canonical wiki example of acknowledging the at-least-once + no-ordering coupling that DLQ-requeue designs introduce.
-
sources/2026-09-08-databricks-build-durable-agents-with-temporal-and-lakebase — at-least-once as the deliberate contract on a workflow↔database boundary, not a message bus. Databricks' durable underwriting agent runs every Lakebase write as a Temporal Activity, which executes at-least-once: a Lakebase write can commit before the Worker reports Activity completion, so if the connection drops in that gap Temporal has no recorded result and schedules another attempt — both attempts are the same logical write. "Temporal determines when to retry. The Activity determines how the external system handles that retry." The consumer-side idempotency mandate is met with stable identities (
run_id/message_id/tool_call_id/event_id/review_id/decision_id) enforced by Postgres PK/unique constraints, plus guarded upserts (patterns/conditional-write) — a retry against an already-terminal row affects zero rows and raises no error. Also the canonical statement of the "every side-effecting tool needs an equivalent contract" rule: a payment API takes an idempotency key, an email service a caller-supplied message ID, a DB a unique constraint; if the external system offers no dedup, the Activity needs its own record or reconciliation. (Source: sources/2026-09-08-databricks-build-durable-agents-with-temporal-and-lakebase) -
sources/2026-09-09-aws-testing-application-resilience-with-amazon-sqs-and-aws-fault-injection-service — at-least-once made visible through a resilience experiment on SQS. On recovery from an injected outage, redelivery + DLQ redrive produce duplicates (the post: count-based metrics "can include retries and duplicates, so treat them as trend indicators rather than exact unique-message counts"), and in-flight messages return to the queue after the visibility timeout — so every consumer must be idempotent and the recovery check is reconciling attempted sends against messages processed plus DLQ plus producer-side fallback storage. Adds the relevance-vs-preservation nuance: after a long outage some redelivered messages represent intent the client already abandoned — drop stale work deliberately rather than reprocessing it. See patterns/dead-letter-queue.
-
sources/2026-10-01-cloudflare-announcing-cloudflare-k2-serverless-event-streams — at-least-once produced by a lease/ack model on a streaming log. Cloudflare K2 hands each
consumecaller a 5-minute lease on a batch; the client then acks (processed — never redelivered), nacks (failed — redeliver), or extends the lease. Lease-expiry-redelivers, so a consumer that crashes (or whose ack is lost) after processing but before ack will see the batch again — the classic "consumer crashes after processing but before acknowledging" duplicate source, surfaced as the broker's delivery contract rather than as an accident. The idempotency mandate therefore lands on K2 consumers directly. K2 is also explicitly batch-grained, not message-grained — it gives up Queues' per-message retries/DLQ "in exchange for" high-scale batch throughput, so the dedup unit is a batch, not an item.
Related¶
- concepts/eventual-consistency — the when-does-it-arrive guarantee this pairs with; at-least-once is the delivery side, eventual consistency is the visibility side.
- concepts/event-driven-architecture — the aggregate shape.
- dynamodb-streams — an at-least-once source.
- transactional-outbox — relies on at-least-once from its relay substrate.
- sqs-dlq-plus-cron-requeue — the concrete pattern that produces replays beyond the initial retry ladder.
Merged aliases¶
at-most-once-delivery-at-least-once-uploads-for-cost-reduction