Skip to content

Patterns

Design patterns: circuit breaker, saga, CQRS, bulkhead, staged rollout, etc.

76 pages

Most-cited

  • Staged rollout 39 sources — Progressively roll out a change — code, config, feature flag — starting in a limited scope and expanding only if health signals stay green.…
  • Specialized agent decomposition 32 sources — Build per-domain agents (storage, databases, client-side traffic, network, …) that each carry a small, well-scoped toolset,…
  • Central proxy choke point 26 sources — Central proxy choke point is the organisational-scale posture of forcing all AI / LLM / agent traffic in an enterprise through one proxy before it reaches any provider,…
  • Shadow migration (dual-run reconciliation) 23 sources — Shadow migration (a.k.a. dual-run with reconciliation) is the pattern of running the new engine in parallel with the old, feeding both the same inputs, producing both outputs,…
  • AI Gateway provider abstraction 20 sources — AI Gateway provider abstraction is the pattern of routing all application LLM calls through a single proxy endpoint that owns provider / model selection, secret injection,…
  • Cheap approximator with expensive fallback 18 sources — Serve most queries with a fast, low-cost ML approximator; fall back to the slow authoritative solver only when the approximator reports high uncertainty.…
  • Tiered storage to object store 17 sources — Split a stateful-broker system's storage into two tiers: a hot local tier (disk, pagecache) holding the most-recent data, and a cold remote tier (object storage — S3, GCS,…
  • Upstream the fix 17 sources — When a performance / correctness / security issue lives in a shared ecosystem primitive (language engine, standard library, OSS framework),…
  • Tool-surface minimization 16 sources — Tool-surface minimization is the discipline of keeping the number of tools an agent sees small, because (a) tool-calling accuracy degrades as the tool inventory grows (arXiv…
  • Progressive configuration rollout 14 sources — Progressive configuration rollout is the same staged-deployment discipline usually applied to code — canary → small cohort → large cohort → fleet,…
  • Fast rollback 13 sources — Ability to revert a change to a known-good state quickly — ideally within seconds — without re-running the full CI/CD pipeline.…
  • Human-calibrated LLM labeling 13 sources — Human-calibrated LLM labeling is the pattern of training a high-volume ML model (a ranker, classifier, preference model, …) on labels generated by an LLM judge that is itself…

All pages (A–Z)

  • Agent sandbox with gateway-only egress — Give the AI agent a real sandbox (container / microVM / OS- level isolation) with full compute, tool, and filesystem access for its reasoning and tool-output post-processing…
  • AI Gateway provider abstraction — AI Gateway provider abstraction is the pattern of routing all application LLM calls through a single proxy endpoint that owns provider / model selection, secret injection,…
  • Alerts as code — Treat each alert as a first-class software artifact: authored with IDE-style tooling, validated against historical data before deploy, diffed on change, and reviewed like code.…
  • Async-projected read model — Async-projected read model is the operational shape of CQRS: a write-optimized source of truth (normalized relational DB, event log, document store) is the command side,…
  • Async replication for cross-region, semi-sync within region — Configure semi-synchronous replication only between replicas within the same region (where cross-AZ latency is single-digit milliseconds) and asynchronous replication…
  • Batch over network to broker — On the producer side of a messaging system, group many small records into one protocol batch before dispatching across the network.…
  • Broker-native Iceberg catalog registration — Broker-native Iceberg catalog registration is the pattern where a streaming broker (writing data as Apache Iceberg snapshots) owns the full lifecycle of its Iceberg tables against…
  • Caching proxy tier — Interpose a stateless proxy tier speaking the cache's native wire protocol between applications and the underlying cache fleet (Redis, Memcached, etc.),…
  • Cell-based architecture for blast-radius reduction — Partition the service into independent cells (self-contained deployable units with isolated data, compute, and control paths) so that any fault's impact is bounded to at most one…
  • Central proxy choke point — Central proxy choke point is the organisational-scale posture of forcing all AI / LLM / agent traffic in an enterprise through one proxy before it reaches any provider,…
  • Cheap approximator with expensive fallback — Serve most queries with a fast, low-cost ML approximator; fall back to the slow authoritative solver only when the approximator reports high uncertainty.…
  • Circuit breaker — A circuit breaker wraps a call to a failure-prone dependency with a state machine that trips open when the dependency is failing at a threshold rate,…
  • Client-proximal leader pinning — On a multi-region cluster, pin each topic's partition leaders to the region where that topic's producer/consumer clients are concentrated,…
  • Closed-loop remediation — Closed-loop remediation couples a detection system directly to an automated action so that a matched finding triggers a corrective response without a human in the loop per…
  • Co-located inference in the serving layer — Co-located inference in the serving layer runs the ML scoring/ranking model inside the same process (or node) that holds the data being scored…
  • Conditional Write (Compare-and-Set on Storage) — A conditional write is a storage-layer primitive that performs a write (typically PUT) only if a precondition on the current state of the target is satisfied…
  • Context-segregated sub-agents — When an agent needs to do work that (a) would consume large amounts of the main context window, (b) needs a different tool surface than the parent,…
  • Coordinator-fronted sharded search — For large search indices that don't fit in a single primary-replica cluster, run a coordinator service in front of multiple per-shard clusters.…
  • Custom benchmarking harness — When a vendor-supplied benchmark tool doesn't match your workload shape or reports the wrong latency metric, write a narrowly-scoped custom harness against the same API.…
  • Custom data structure for hot path — When stdlib or well-known crate data structures don't match your workload shape on a genuine hot path, write one that does.…
  • Data contract — A data contract is a formal, explicit agreement between the producer of a data product and its consumers that specifies what the product is, how it behaves, who may use it,…
  • Database branch per test over mocking — On a substrate where database branching is sub-second and cheap (copy-on-write storage fork; e.g. Lakebase / Neon), replace database-interface mocks with per-test / per-PR / per-…
  • Dead-letter queue — A dead-letter queue (DLQ) is a secondary queue that receives messages a consumer repeatedly fails to process, so they are set aside for investigation instead of blocking…
  • Default-on security upgrade at no additional cost — A product-strategy pattern where an infrastructure provider ships a security capability as a universal, default-enabled platform behaviour…
  • Deterministic tool vs LLM judgment — Deterministic tool vs LLM judgment is the discipline of deciding, for each unit of work in an agent system, whether it needs an LLM or a deterministic function…
  • Dispatcher–Coding-Agent–Closer — A three-part loop for putting recurring, predictable engineering maintenance work (KTLO — "keep the lights on") on autopilot with a coding agent,…
  • Disposable VM for agentic loop — Instead of running an LLM-driven agentic coding loop on the developer's laptop (or on any shared dev server), spin up a disposable, clean-slate VM per task / session,…
  • Durable event log as agent audit envelope — Every agent interaction — prompt, input, context retrieval, tool call, output, and action — is captured as a first-class durable event on a streaming log.…
  • Experimental-tier model promotion — Experimental-tier model promotion is the lifecycle for adopting a third-party frontier model of unknown quality into a large user population: give everyone access immediately…
  • Fast rollback — Ability to revert a change to a known-good state quickly — ideally within seconds — without re-running the full CI/CD pipeline.…
  • Golden path with escapes — Golden path with escapes is the platform-team design where the default service-creation / service-config flow is heavily opinionated — consistent defaults,…
  • Hedged reads for tail latency — Issue redundant read requests to multiple storage nodes; use the first response and discard the rest, mitigating single-node latency outliers (laggards).
  • Human-calibrated LLM labeling — Human-calibrated LLM labeling is the pattern of training a high-volume ML model (a ranker, classifier, preference model, …) on labels generated by an LLM judge that is itself…
  • LTX compaction (time-window merge of SQLite page runs) — Represent a database's changes as sorted per-transaction page-range files (LTX), then periodically k-way-merge adjacent time windows into larger files that keep only the latest…
  • MCP as centralized integration proxy — Deploy a single MCP server tier in front of the enterprise's internal systems (databases, queues, SaaS APIs, code repos, docs) and make it the mandatory choke-point through which…
  • Measurement-driven micro-optimization — Pick the code worth optimizing by production profiling, not by taste; validate each candidate change against a repeatable benchmark; ship;…
  • Multimodal content understanding — Multimodal content understanding is the ingestion-time pattern of routing each content type to its own specialized extraction path — documents, images, PDFs, audio,…
  • On-behalf-of (OBO) agent authorization — When an AI agent makes a tool call on behalf of an authenticated human user (or a calling service), the tool-invocation boundary (typically an MCP server) forwards the call…
  • Parallel retrieval fusion — No single retrieval method works best for all queries. The distribution of query shapes looks like:
  • Partner managed service as native binding — Integrate a third-party managed service (database, vector store, inference provider) into a platform such that customer code consumes it through the same binding mechanism…
  • Per-developer database branch paired with code branch — When a developer creates a git feature branch, automatically provision a matching database branch off production (or a golden baseline).…
  • Pilot light deployment — DR deployment tier where the data tier in the secondary environment is running and replicated, but the compute tier is stopped (or minimally provisioned).…
  • Power of Two Choices (P2C) — Power of Two Choices (P2C): instead of picking one backend uniformly at random, pick two at random and route the request to the one with fewer active requests / lower observed…
  • Presentation Layer Over Storage — Presentation-layer-over-storage treats an application-facing data interface (filesystem, SQL table, vector index, message queue) as a presentation of canonical data that physically…
  • Progressive configuration rollout — Progressive configuration rollout is the same staged-deployment discipline usually applied to code — canary → small cohort → large cohort → fleet,…
  • Prompt optimizer flywheel — Prompt optimizer flywheel is the pattern of closing a feedback loop between an LLM judge, a structured representation of judge-vs-human disagreements, and a prompt optimizer (e.g.…
  • Property-Based Testing — Property-based testing is a testing pattern where, instead of asserting outputs on specific inputs (traditional example-based unit tests), you:
  • Protocol algorithm negotiation — Protocol algorithm negotiation is the protocol-design pattern where each side of a connection advertises its supported algorithms (by name, in preference order),…
  • Prototype before production — Prototype before production — before committing to an architectural choice that will be expensive to revisit, build a standalone simulator that exposes the choice space cheaply…
  • Quarterly Internet disruption review — Pattern: On a recurring quarterly cadence, an Internet-observability team (one that owns a large-scale traffic vantage point — CDN, reverse proxy, DNS resolver,…
  • Saga over long-running transaction — Decompose a logically-atomic multi-step workflow from a single long-running database transaction into a sequence of short local transactions connected by compensating actions,…
  • Separate revoke from establish in leader election — Traditional majority-quorum consensus algorithms (Paxos, Raft) perform revocation of the previous leader and establishment of the new leader as a single atomic action…
  • Shadow migration (dual-run reconciliation) — Shadow migration (a.k.a. dual-run with reconciliation) is the pattern of running the new engine in parallel with the old, feeding both the same inputs, producing both outputs,…
  • Shard key aligned with query pattern — Choose the shard key so that the dominant query's predicate contains the shard key. The single-most-common query routes to exactly one shard;…
  • Signed Bot Request (Ed25519 + JWK directory + RFC 9421) — Signed Bot Request is the design pattern for giving an automated client (crawler / bot / agent) a cryptographic identity that an origin can verify per-request,…
  • Snapshot plus catch-up replication — Copying a live database's data to another system without taking the source offline has a fundamental tension:
  • Snapshot-replay agent evaluation — Capture snapshots of production-state inputs (queries, tool responses, intermediate state) from real agent runs, then replay them through candidate agent configurations (new…
  • Specialized agent decomposition — Build per-domain agents (storage, databases, client-side traffic, network, …) that each carry a small, well-scoped toolset,…
  • SQLite + LiteFS + Litestream — Use SQLite as the primary storage engine, LiteFS as the distributed primary/replica filesystem layer (subsecond replication + primary failover),…
  • Staged rollout — Progressively roll out a change — code, config, feature flag — starting in a limited scope and expanding only if health signals stay green.…
  • Strangler Fig — Strangler Fig (Martin Fowler, 2004) is the pattern of incrementally replacing a legacy system by routing requests through an interception layer — a proxy, gateway,…
  • Stratified evaluation sampling — Stratified evaluation sampling is the production-loop pattern of scoring only a fraction of live traffic, with per-stratum sampling rates weighted by risk and value instead…
  • Streaming broker as lakehouse Bronze sink — Problem. Most organisations running a Kafka-class streaming broker for operational data also run a lakehouse with a Medallion Architecture for analytics.…
  • Teacher-Student Model Compression — Teacher-student model compression is the engineering pattern of wrapping knowledge distillation into a production deployment shape: pick a model class that solves the task…
  • Telemetry to Lakehouse — Telemetry to Lakehouse is the pattern of landing operational / tool / agent telemetry directly into governed open-table-format tables (typically Delta Lake or Iceberg) instead…
  • Tests as executable specifications — Treat the test suite not just as a regression net, but as the behavioral specification of the system — a corpus of executable assertions that both human reviewers and AI agents…
  • Tiered storage to object store — Split a stateful-broker system's storage into two tiers: a hot local tier (disk, pagecache) holding the most-recent data, and a cold remote tier (object storage — S3, GCS,…
  • Tool-surface minimization — Tool-surface minimization is the discipline of keeping the number of tools an agent sees small, because (a) tool-calling accuracy degrades as the tool inventory grows (arXiv…
  • TTL-based resource lifecycle — A resource management pattern where every provisioned environment receives a default time-to-live (TTL) with extension options, automatic expiration notifications,…
  • Two-stage evaluation — Two-stage evaluation is the pattern of splitting a per-event match/decision pipeline into:
  • Unified billing across providers — Unified billing across providers is the cost-management pattern of routing all LLM / AI traffic — first-party inference capacity, BYO external provider keys,…
  • Upstream contribution parallel to in-house integration — When you need a new capability that doesn't exist in an upstream open-source project, and you cannot afford to wait on the upstream maintainer's merge timeline,…
  • Upstream the fix — When a performance / correctness / security issue lives in a shared ecosystem primitive (language engine, standard library, OSS framework),…
  • Verify before changing code — When an automated coding agent acts on a ticket, treat the ticket as context, not as proof of the current state of the code.…
  • Warm pool, zero-work create path — A VM / compute primitive has a user-visible create operation, and the DX requires it to feel instantaneous — sub-2-second, ideally sub-second.…
  • Wrap CLI as MCP server — Expose an existing CLI as an LLM tool surface by writing a thin MCP server that:
Last updated · 766 distilled / 2,225 read