Skip to content

SYSTEM Cited by 7 sources

AWS SQS

Amazon SQS (Simple Queue Service) is AWS's managed message queue: durable, at-least-once delivery, essentially unlimited scale, with standard and FIFO variants. In system-design terms, it's the generic durable work-dispatch primitive between producers and consumers that can operate at different rates and need the queue to survive either side crashing.

Typical role in data pipelines

SQS commonly appears as the durability layer in warehouse-unload bridges: the warehouse exports into S3, an S3-event emits an SQS message, and the ingester consumes the queue and writes to the serving store. If the ingester crashes or the serving store throttles, the message stays in the queue; when the ingester comes back up, it resumes without data loss.

Canva uses exactly this shape in the counting pipeline: Snowflake unload → S3 → SQS → rate-limited ingester → service RDS. They call out SQS's durability specifically as the reason exported data "doesn't get lost". (Source: sources/2024-04-29-canva-scaling-to-count-billions; warehouse-unload-bridge)

Operational notes

  • At-least-once means consumers must be idempotent (outer-join upserts are one way to get there — see end-to-end-recompute).
  • Visibility timeout has to be tuned against downstream write latency; too low and messages double-deliver, too high and failed messages sit idle.
  • DLQs for poison-pill messages are a must in any serious pipeline.

Seen in

  • sources/2026-09-10-aws-building-resilient-real-time-streaming-workers-with-amazon-dynamodb-leases — SQS as a fast-notification hint, not the source of ownership. In AWS's WebSocket-fleet pattern a START event enqueues an SQS message so workers claim new connections in seconds instead of waiting for the next 60 s reconciliation cycle — but the message only hints; the worker must still win the DynamoDB lease before opening the connection. The post articulates the general boundary: SQS delivers a unit of work to one consumer and has no notion of continuous ownership — it cannot track who currently owns a connection, query for unowned ones, or hold domain state — so it complements, but cannot replace, a lease. Because delivery is at-least-once, the lease-before-connect check also dedups redeliveries. (Source: sources/2026-09-10-aws-building-resilient-real-time-streaming-workers-with-amazon-dynamodb-leases)
  • sources/2024-04-29-canva-scaling-to-count-billions — SQS between S3 and the RDS ingester in Canva's warehouse-unload bridge, providing durability for warehouse→OLTP export.
  • sources/2024-07-29-aws-amazons-exabyte-scale-migration-from-apache-spark-to-ray-on-ec2 — SQS as one of the primitives in Amazon Retail BDT's 2021 serverless Ray job-management substrate (alongside systems/dynamodb, systems/aws-sns, systems/aws-s3) for durable job-lifecycle tracking across thousands of exabyte-scale Ray compaction jobs per day.
  • sources/2026-02-04-aws-amazon-key-eventbridge-event-driven-architecture — Named (with SNS) in the "ad-hoc SNS/SQS pairs" anti-pattern that Amazon Key's EventBridge migration replaced. SQS still the natural per-subscriber queue under an EventBridge target, just not the shared-bus abstraction.
  • — SQS as the attached dead-letter queue of a Lambda outbox relay. Zalando Payments's Order Store publishes events to Nakadi via a Lambda triggered by DynamoDB Streams; when the Lambda's exponential-backoff retries exhaust, the event lands in the Lambda's built-in SQS DLQ. A Kubernetes CronJob runs the same Python publication code on an interval, draining the DLQ until Nakadi accepts — the same-code-on-two-substrates property is load-bearing. Canonical wiki example of SQS DLQ + external cron re-drain as the fallback tier of an event-publish pipeline. See sqs-dlq-plus-cron-requeue, dynamodb-streams-plus-lambda-outbox-relay.
  • sources/2026-05-04-netflix-democratizing-machine-learning-building-the-model-lifecycle-graph — SQS (paired with SNS + Kafka) as the per-subscriber durability layer for Netflix MDS's ingestion of thin notification-of-change events from six source systems. SNS fans out source-system events; SQS provides per-consumer durability so MDS can absorb ingestion bursts without dropping events while the enrichment workers are hydrating from source APIs at a rate-limited cadence.

  • sources/2026-09-09-aws-testing-application-resilience-with-amazon-sqs-and-aws-fault-injection-service — SQS as the target of a resilience experiment, and the canonical disclosure of its DLQ mechanics under failure. An FIS-driven scoped deny blocks data-plane ops (SendMessage/ReceiveMessage/DeleteMessage/ ChangeMessageVisibility/PurgeQueue). Key operational facts surfaced: (1) redrive is driven by maxReceiveCount — a message dead-letters only after being received that many times, so with ReceiveMessage denied the DLQ stays empty during the outage and only moves on recovery; (2) a DLQ has no depth limit — the constraint is the retention period (default 4 days, max 14), and for standard queues the retention clock runs from original enqueue time and does not reset on the move to the DLQ (FIFO does reset); (3) visibility timeout returns in-flight messages to the queue when a consumer stalls; (4) count-based metrics include retries/duplicates. See patterns/dead-letter-queue.

  • sources/2026-09-30-aws-how-mhk-built-a-hipaa-eligible-agentic-ai-solution-on-amazon-bedrock — SQS as the event-driven backbone of a controller-agent agentic workflow. In MHK's SmartProminence AI Orchestrator the orchestration core enqueues a ControllerTaskMessage on the Controller Invoke Queue to start a workflow and an AgentTaskMessage on the Agent Invoke Queue to dispatch a step; the only blocking call in the whole pipeline is the Bedrock LLM invocation — everything else is async over SQS, letting the system run "hundreds of concurrent jobs without resource contention." Each message carries a per-dispatch, KMS-signed capability token scoped to only that operation's data (so a compromised consumer can't reach other steps/clients), and dead-letter queues capture failed tasks so a failure doesn't block parallel steps. Canonical wiki example of SQS as the decoupling layer in a stateless, independently-scalable event-driven agent-orchestration system. (Source: sources/2026-09-30-aws-how-mhk-built-a-hipaa-eligible-agentic-ai-solution-on-amazon-bedrock)
Last updated · 766 distilled / 2,225 read