Skip to content

Concepts

Distributed-systems concepts: consistency models, CAP, consensus, backpressure, CRDTs, etc.

282 pages

Most-cited

  • Blast radius 48 sources — Blast radius is the scope of damage that a single fault — bug, misconfiguration, vulnerability, runaway workload, compromised credential…
  • Observability 46 sources — The function of providing visibility into application performance and reliability via metrics, logs, and traces. The core operational quality it serves: lowering MTTD (mean time…
  • Control plane / data plane separation 44 sources — Architectural split between the "decide" path (control plane: validation, authorization, policy, rollout decisions, scheduling) and the "deliver" path (data plane: storage,…
  • Compute–storage separation 31 sources — Compute–storage separation is the architectural property where a system's persistence layer and its query/compute layer are decoupled and scale independently.…
  • Context engineering 30 sources — Context engineering is the discipline of allocating a fixed token budget across the components that compete for the LLM's context window — system prompts, tool descriptions,…
  • Change Data Capture (CDC) 25 sources — Change Data Capture (CDC) is the practice of materialising a table as an ongoing stream of insert / update / delete deltas rather than (or alongside) an authoritative "current…
  • Defense in Depth 25 sources — Defense in depth is the security posture of stacking independent protective layers such that no single compromised or misconfigured layer exposes the whole system.…
  • Scale to Zero 24 sources — Scale-to-zero is a service-design property in which an application consumes no capacity and accrues no charge when it has no traffic,…
  • Tenant isolation 24 sources — Tenant isolation in a multi-tenant SaaS is the property that one tenant's users, data, and policies are strictly inaccessible to other tenants…
  • Cache hit rate 22 sources — The cache hit rate is the fraction of data requests served from the cache without needing to go to the slower backing store:
  • LLM as Judge 22 sources — LLM-as-judge is the evaluation pattern in which one LLM scores another model's (or agent's) output against a rubric — accuracy, helpfulness, policy adherence,…
  • Vector Similarity Search 22 sources — Vector similarity search is the retrieval primitive behind semantic search, recommendation, and RAG: given a query vector and a corpus of vectors (usually embeddings of documents /…

All pages (A–Z)

  • ACID properties — ACID is a set of four properties — Atomicity, Consistency, Isolation, Durability — that together define the guarantees a database transaction offers.…
  • Actor model — The actor model is a concurrency + distributed-systems programming model in which:
  • Agent-ergonomic CLI — An agent-ergonomic CLI is a command-line interface explicitly designed for AI agents as the primary caller, with human usability as a correlated-but-secondary concern.…
  • Agent memory — Agent memory is an AI agent's accumulated, searchable context across turns and sessions — the things this agent (or this agent for this user) has seen, decided,…
  • Agentic development loop — The agentic development loop is a closed-loop LLM code generation workflow: the LLM proposes code, an execution environment runs the code, the environment's output (stdout,…
  • AI agent guardrails — AI agent guardrails is the discipline of running AI-generated code through the same (or stronger) quality gates that human-written code would face,…
  • Alert fatigue — The operator-side failure mode in which a notification channel emits so many alerts — especially low-signal / duplicate / flapping…
  • Anomaly detection — Anomaly detection is the automated identification of observations that deviate from a system's learned "normal" behavior…
  • Anycast — Anycast is a network-layer routing discipline in which the same IP address (or prefix) is advertised from multiple geographically dispersed points of presence (POPs) via BGP,…
  • Assignment problem — An assignment problem is the combinatorial-optimization problem of assigning a set of objects to a set of bins so as to optimize one or more objectives while satisfying…
  • Asynchronous replication — Asynchronous replication is a replication posture in which the primary (source) system acknowledges a write to the client before secondary (replica) systems have applied it.…
  • At-least-once delivery — At-least-once delivery is the messaging guarantee that each message will eventually be delivered to a consumer at least once,…
  • Attack surface minimization — Attack surface minimization is the design discipline of keeping the set of code paths / APIs / parsers / features reachable by untrusted input as small as possible.…
  • Attribute-based access control (ABAC) — Attribute-based access control (ABAC) decides whether a principal may perform an action on a resource by evaluating a policy over attributes of all three (plus context)…
  • Audit trail — An audit trail is a durable, queryable record of every state-changing operation performed against a system — who changed what, when,…
  • Autonomous System (AS) — An Autonomous System is an administratively-unified network that speaks BGP with its neighbors under a single routing policy.…
  • B-tree — A B-tree of order K is a self-balancing tree where each node stores between 1 and K key/value pairs (internal nodes ≥K/2), every node has N+1 children for N keys,…
  • Back-of-the-envelope estimation — Back-of-the-envelope estimation (BOTE) is the engineering discipline of arriving at a rough-order-of-magnitude sizing number for a proposed system — capacity, cost, latency,…
  • Backpressure — Backpressure is the control-plane primitive by which a slow consumer in a streaming pipeline signals a fast producer to slow down.…
  • Backward compatibility — Backward compatibility is the property that interfaces accept and behave correctly on inputs + requests that were valid under prior versions of the interface.…
  • Batching latency trade-off — The batching latency trade-off is the explicit exchange a producer makes when it groups records into batches: higher throughput is paid for with higher per-record latency.…
  • Benchmark methodology bias — A benchmark methodology bias is a confounder built into a benchmark's setup (not its subject) that systematically skews results in one direction — and critically,…
  • Border Gateway Protocol (BGP) — BGP is the path-vector routing protocol that glues the Internet together. Each Autonomous System (AS) speaks BGP with its neighbors to exchange reachability information…
  • Bin-packing — Bin-packing is the combinatorial-optimisation problem of placing a set of items with known sizes into the smallest number of fixed-capacity bins (or equivalently: maximising…
  • Binary-size bloat — Binary-size bloat is the monotonic growth of a compiled artifact over time as features + dependencies accumulate, without a commensurate removal discipline.…
  • Blast radius — Blast radius is the scope of damage that a single fault — bug, misconfiguration, vulnerability, runaway workload, compromised credential…
  • Bloom filter — A Bloom filter is a space-efficient probabilistic data structure that tests whether an element is a member of a set. It returns:
  • BYOK (Bring Your Own Key) — BYOK (Bring-Your-Own-Key) is the posture in which a customer stores a third-party provider API key (OpenAI, Anthropic, Google,…
  • Cache hit rate — The cache hit rate is the fraction of data requests served from the cache without needing to go to the slower backing store:
  • Cache locality — Cache locality is the property that requests for the same key consistently arrive at the same node, so a node-local cache keyed by that key accumulates hits across those requests…
  • Cache-variant explosion — Cache-variant explosion (a.k.a. cache fragmentation) is the failure mode in which a single cacheable resource is split into many cache entries — one per distinct cache-key value…
  • Capability-based sandbox — A capability-based sandbox is a code-execution environment that starts with no ambient authority — no network access, no filesystem access, no secrets,…
  • Cascading failure — A failure that grows over time via a positive feedback loop: one node's overload spreads load to remaining nodes, increasing their probability of failure, which shifts more load,…
  • Cell-based architecture — Cell-based architecture is the design pattern of partitioning a service into multiple independent deployable units ("cells") — each with its own compute, storage,…
  • Centralized AI governance — Centralized AI governance is the pattern-concept of routing all of an organisation's AI traffic (LLM calls + tool calls + MCP traffic + coding-agent activity) through one policy…
  • Change Data Capture (CDC) — Change Data Capture (CDC) is the practice of materialising a table as an ongoing stream of insert / update / delete deltas rather than (or alongside) an authoritative "current…
  • Chaos engineering — Chaos engineering is the discipline of continuously inducing controlled failures in a production system to verify that its fault-tolerance design actually works.…
  • Circular dependency (deployment context) — A circular dependency in the deployment context is the failure mode where the act of deploying a fix for a service depends — directly or indirectly…
  • Client-server model — The client-server model is the foundational deployment pattern of the Internet: a client makes a request to a server, which responds with a resource.…
  • Client-side load balancing — Client-side load balancing means the caller chooses which backend instance to send a request to, rather than delegating that decision to a proxy (L7 sidecar,…
  • Cold Start — Cold start names two distinct phenomena that share a name because both describe "the first time around is slow / hard":
  • Columnar storage format — A columnar storage format lays out a tabular dataset on disk by column rather than by row: all values of column A are contiguous, then all values of column B, etc.…
  • Compression codec trade-off — Compression codec trade-off is the choice a streaming producer makes between space / bandwidth savings (high compression ratio = fewer bytes over the wire, less disk,…
  • Compute–storage separation — Compute–storage separation is the architectural property where a system's persistence layer and its query/compute layer are decoupled and scale independently.…
  • Confidential computing — Confidential computing is the posture of protecting data in use — i.e. plaintext that is being actively computed on — via hardware-enforced isolation primitives (TEEs).…
  • Confidential storage inside the TEE — Confidential storage inside the TEE is the architectural move of placing the storage engine — not just the compute — inside the TEE trust boundary,…
  • Conflict-free Replicated Data Type (CRDT) — A Conflict-free Replicated Data Type (CRDT) is a data structure replicated across multiple machines such that: (1) any replica may be updated independently and concurrently without…
  • Confused deputy problem — The confused deputy problem occurs when a trusted intermediary (the "deputy") is tricked into misusing its legitimate authority on behalf of an attacker who lacks that authority…
  • Connection pool exhaustion — A connection pool is a fixed-size set of pre-established database connections shared across application workers. An application worker takes a connection from the pool…
  • Consensus algorithm — A consensus algorithm allows a set of machines communicating over a network to agree on the same sequence of values (e.g.,…
  • Consistent hashing — Consistent hashing maps keys to buckets (shards, servers, cohorts) in a way that minimises re-mapping when the bucket set changes.…
  • Content negotiation — Content negotiation is the HTTP mechanism by which a single URL can return more than one correct representation, with the server selecting among them based on request headers.…
  • Context compaction — Context compaction is the practice of trimming or summarising older content in an LLM agent's context window to keep a long-running reasoning loop within token limits without…
  • Context engineering — Context engineering is the discipline of allocating a fixed token budget across the components that compete for the LLM's context window — system prompts, tool descriptions,…
  • Context switch — A context switch is the OS kernel's act of saving the current process (or thread)'s execution state and restoring another one's, so the CPU can resume a different unit of work.…
  • Context window as token budget — The context window supplied to an LLM call is a fixed token budget. Every input the program keeps in that window — user messages, assistant replies,…
  • Control plane / data plane separation — Architectural split between the "decide" path (control plane: validation, authorization, policy, rollout decisions, scheduling) and the "deliver" path (data plane: storage,…
  • Conway's Law — Conway's Law (Melvin Conway, 1968) is the sociotechnical observation that systems tend to mirror the communication structure of the organizations that build them.…
  • Coordinated disclosure — Coordinated disclosure (historically responsible disclosure) is the industry norm by which a security vulnerability is not made public until the affected vendor has had…
  • Copy-on-write storage fork — A copy-on-write storage fork is a storage-cloning mechanism that creates a second logical copy of a dataset without initially duplicating the underlying pages…
  • Core Web Vitals — Core Web Vitals is Google's standardised set of user- experience metrics for web pages, measured in the browser and intended as proxies for perceived page quality (loading,…
  • CQRS (Command-Query Responsibility Segregation) — CQRS — Command-Query Responsibility Segregation — is the idea that the data model optimized for accepting writes (commands) doesn't have to be the same data model serving reads…
  • CRDT (conflict-free replicated data type) — A Conflict-free Replicated Data Type is a data structure whose replicas can be updated concurrently and without coordination and will still converge to the same value…
  • Critical path (build / pipeline / distributed DAG) — In a DAG of dependent actions (build targets, pipeline steps, tasks), the critical path is the longest chain of dependent work from the graph's root to its sink.…
  • Data Lakehouse — A data lakehouse is a data-platform architectural class that combines:
  • Data lineage — Data lineage is the graph of relationships between data assets that tracks "where did this data come from" and "where does this data flow to" — source → sink relationships…
  • Data Mesh — Data mesh is an architectural approach to organisation-wide data sharing in which data products are owned and exposed by domain teams (R&D, After-Sales, Marketing, etc.),…
  • Data ontology — A data ontology is a semantic context layer that captures what data means in the context of a business — its definitions, relationships, calculations, authoritative sources,…
  • Data residency — Data residency constrains the legal or contractual geography in which business data may be stored, processed, recovered, or made accessible.…
  • Database branching — Database branching is the workflow primitive of creating an isolated sandbox copy of a production database's schema (and, on some platforms,…
  • Defense in Depth — Defense in depth is the security posture of stacking independent protective layers such that no single compromised or misconfigured layer exposes the whole system.…
  • Deterministic Simulation — Deterministic simulation is a testing discipline where the entire system under test runs inside a custom executor/scheduler that eliminates every source of non-determinism it…
  • Differential privacy — Differential privacy (DP) is a mathematical guarantee that the output of a computation is statistically insensitive to whether any single individual's data was included or not.…
  • Digital sovereignty — Digital sovereignty is "managing digital dependencies — deciding how data, technologies, and infrastructure are used, and reducing the risk of loss of access, control,…
  • Distributed lease — A distributed lease is a time-bounded ownership claim over a resource (a connection, a partition, a shard, a "leader" role, a file) that a process must continuously renew to keep.…
  • Distributed monolith — A distributed monolith is the anti-pattern of having successfully decomposed a monolith into many microservices — but keeping the underlying coupling intact.…
  • Distributed transactions — Distributed transactions are atomic operations that span multiple database rows, shards, or nodes — committing all parts or none, across machine boundaries.…
  • Double-checked locking — Double-checked locking is a concurrency idiom that avoids duplicating expensive work when multiple threads / requests / processes compete to do the same thing. The structure is:
  • Downgrade attack — A downgrade attack is an active-adversary manipulation of a cryptographic protocol negotiation that forces the two endpoints to select a weaker algorithm (or a lower protocol…
  • Durable execution — Durable execution is the property of a long-running computation (multi-minute LLM loops, multi-hour CI pipelines, multi-day workflows) that it survives any interruption of its host…
  • Efficiency frontier — The efficiency frontier is the set of models that offer the best price point for a given level of intelligence. It is deliberately contrasted with the intelligence frontier…
  • Egress Cost — Egress cost is the per-byte charge a cloud provider levies when data leaves a region, a cloud, or (in some cases) an availability zone.…
  • Elasticity — Elasticity is the property that a service's capacity and performance expand and contract to customer demand without requiring the customer to forecast, provision,…
  • ELT vs ETL — ETL (Extract → Transform → Load) transforms data before loading it into the warehouse, usually in an external worker. ELT (Extract → Load → Transform) lands raw data into…
  • End-to-end encryption (E2EE) — End-to-end encryption (E2EE) is the property that a message is encrypted on the sender's device, decrypted only on the recipient's device,…
  • Entity resolution — Entity resolution (ER) — also called record linkage, deduplication, or identity resolution — is the problem of deciding whether two or more records,…
  • Envelope Encryption — Envelope encryption is a multi-level key-hierarchy scheme for encrypting data at rest. Data is encrypted with unique per-segment Data Encryption Keys (DEKs),…
  • Ephemeral credentials — Ephemeral credentials are credentials — keys, tokens, certificates, or passwords — that are generated on demand, used for a short lifetime, and discarded.…
  • Erasure coding — Erasure coding is a redundancy scheme that encodes data into more pieces than are needed to read it, so that the data survives the loss of any fixed number of pieces.…
  • Error budget — The error budget is the complement of an SLO: if the SLO target is 99.9% over a 28-day window, the error budget is the 0.1% of traffic the service is allowed to fail.…
  • Event-driven architecture — Event-driven architecture (EDA) is a software-system style in which services communicate asynchronously by publishing and subscribing to events on a shared bus…
  • Eventual consistency — Eventual consistency is a liveness guarantee: if no new updates are made to a shared value, all observers will eventually converge on the same read.…
  • Evolutionary database design — Evolutionary database design is the discipline of treating database schemas as first-class artefacts that evolve incrementally alongside application code,…
  • Exponential backoff with jitter — Exponential backoff with jitter is the retry-scheduling strategy that pairs two disciplines:
  • Fail-open vs fail-closed — A design choice for what a module does when its input is corrupt, out-of-range, or fails an invariant:
  • Fair sharing — Fair sharing is a resource allocation policy where compute capacity is distributed among tenants proportionally to their configured weights,…
  • Fast VM boot DX — The developer-experience property that a VM primitive can be treated like a container or a function: started on demand per request / per call / per session,…
  • Feature flag — A feature flag is a runtime switch that gates a code path by context, rollout percentage, or targeting rule — enabling release to be decoupled from deploy and cohort-scoped…
  • Feature store — A feature store is the class of ML-infrastructure systems that manage and deliver feature data — the numerical/categorical signals a model consumes at both training time…
  • File vs. Object Semantics — File semantics (the OS filesystem contract applications have been written against for 50 years) and object semantics (the S3-style immutable-blob-with-HTTP-API contract) differ…
  • Fine-grained authorization — Fine-grained authorization means deciding access at the level of individual resources, actions, and runtime context — not just at the level of roles or API endpoints.…
  • Flamegraph profiling — Flamegraph profiling is the practice of sampling a running process's stack at high frequency and rendering the aggregated stacks as a flamegraph — a horizontal-axis-is-samples,…
  • Forward security — Forward security (historically forward secrecy in the TLS literature) is the property that compromise of long-term key material at time T does not expose data from sessions…
  • Garbage collection (storage) — In immutable / append-only storage systems, garbage collection (GC) is the stage that identifies which blobs, rows, or objects are no longer referenced and marks them safe…
  • Geographic sharding — Geographic sharding partitions data by a location dimension — usually the user's destination country / region / continent…
  • Goodput — Goodput in large-scale GPU training is the proportion of GPU time spent on productive computation rather than waiting on inputs or recovering from failures.…
  • Gossip protocol — A gossip protocol (also called an epidemic protocol — the transmission of messages is analogous to how epidemics spread, or how rumours spread in a crowd) is a family…
  • Governed agent data access — Governed agent data access is a two-axis design surface — access controls (which agent gets what data, on whose behalf, under what consent) + observability (what did the agent…
  • Graceful degradation — Graceful degradation is the design property that when a subsystem or dependency fails, the service as a whole continues to operate in a reduced-but-useful mode,…
  • Grey failure — Grey failure names a component that is not fully broken but not fully healthy — partially, intermittently, or sub-specification degraded.…
  • Hard-drive physics (capacity vs. seek-time) — Hard drives are mechanical devices (spinning platters + moving arm + flying head). This constrains them in a way that grows more severe every generation: capacity scales fast,…
  • Harvest-now, decrypt-later (HNDL) — Harvest-now, decrypt-later (also store-now-decrypt-later / steal-now-decrypt-later) is the threat model in which an adversary captures encrypted traffic today and stores it…
  • Head-of-line blocking — Head-of-line (HOL) blocking occurs when a slow item at the front of an ordered queue or stream prevents subsequent items from being processed,…
  • Hexagonal architecture — Hexagonal architecture (Alistair Cockburn, 2005), also known as ports-and-adapters, organizes a codebase into concentric layers where the innermost domain holds business rules…
  • Horizontal sharding — Horizontal sharding splits a single logical table (or group of related tables) so that its rows live across multiple physical database instances.…
  • Hot key — A hot key is a single key whose request rate is disproportionately higher than the rest of the keyspace. Under any scheme that maps one key to one node (horizontal-sharding /…
  • Hot path — The hot path is the code that runs on every (or near-every) request, especially at high request rates. It's the call chain whose per-invocation cost is multiplied by the request…
  • Human-in-the-loop — Human-in-the-loop (HITL) is an architectural stance in which an automated or AI system prepares, ranks, and proposes decisions but a human retains final authority over…
  • Hybrid Search — Hybrid search combines lexical retrieval (BM25 / keyword / exact-term) with semantic retrieval (dense vector similarity)…
  • Iceberg topic — An Iceberg topic is a streaming-broker construct — introduced by Redpanda in 2024 — in which a single logical entity is both a Kafka-protocol topic and an Apache Iceberg table…
  • Idempotent operations — An idempotent operation is one whose effect is the same whether executed once or many times with the same inputs. Idempotency is the enabling property for safe retries: without it…
  • Immutable Object Storage — Immutable object storage is a model in which the stored unit — the object — cannot be partially modified after it is written.…
  • In-kernel filtering — In-kernel filtering is the architectural move of evaluating per-event match/drop decisions inside the kernel (via ebpf), before events are handed off to user-space consumers.…
  • Incremental View Maintenance (IVM) — Incremental view maintenance (IVM) is the database technique of keeping a materialized view up to date by applying only the deltas induced by changes to its inputs,…
  • Integration tests against real database — Integration tests against real database is the testing discipline in which a test case runs against a live instance of the actual database technology that production uses…
  • IO wait — %iowait on Linux is the fraction of CPU time reported as idle while at least one CPU-local task is blocked on disk I/O. It appears as the wa column in vmstat and top,…
  • IVF (Inverted File Index) — IVF (Inverted File index) is a cluster-based approach to approximate nearest-neighbor (ANN) search. The corpus is partitioned into clusters by a coarse quantizer (typically k-means…
  • Knowledge Distillation — Knowledge distillation is the technique of transferring knowledge from a large, capable, expensive-to-run teacher model to a small, fast,…
  • Knowledge graph — A knowledge graph is a data structure that captures relationships between entities (people, documents, events, projects, activities) rather than just their individual contents.…
  • Kubernetes Operator pattern — A Kubernetes Operator is a custom controller that extends the Kubernetes Control Loop (observe → diff → act) to manage a domain- specific workload…
  • KV cache (transformer inference) — The KV cache is the per-layer, per-token Key and Value projection tensor store that a transformer decoder reuses across autoregressive generation steps.…
  • Last-write-wins — Last-write-wins (LWW) is a conflict-resolution rule for concurrent updates to the same state: if two replicas update the same field, the one with the later timestamp wins.
  • Late-arriving data — Late-arriving data is event data that reaches the analytics / streaming system after the "wall clock" bucket that the event belongs to has already started being processed…
  • Latency budget — A latency budget is the total wall-clock time allocated to a request path before the user experience degrades or a business SLO is violated.…
  • Latency-critical vs latency-tolerant workload — Latency-critical vs latency-tolerant workload is the workload- class distinction between streams whose business value depends on low end-to-end latency (tight p99 / p99.9 targets)…
  • Latent misconfiguration — Latent misconfiguration is a configuration bug that is structurally wrong from the moment it lands in production but produces no observable effect until some later,…
  • Leader election — Leader election is a coordination primitive in distributed systems where multiple instances (or nodes) select exactly one among them to perform a particular role or action at any…
  • Leader-follower replication (Kafka partition) — In Apache Kafka, every partition has exactly one leader at a time; the other replicas are followers. Writes always go to the leader;…
  • Least-privileged access — Each actor — user, service, operator — sees only the data and gets only the capabilities strictly required for their current interaction.…
  • Lightweight formal verification — Lightweight formal verification is a family of techniques that sit between ad-hoc testing and heavyweight proof-based formal methods (TLA+, Coq, Isabelle).…
  • Linearizability — Linearizability is the strongest possible consistency level for a distributed data system. Operations are ordered exactly as they occurred in real time…
  • List virtualization — List virtualization (a.k.a. windowing) is a front-end rendering technique for long scrollable lists: mount only the rows currently on screen (plus a small over-scan margin),…
  • Little's Law — Little's Law is the foundational result in queueing theory due to John Little (1961, MIT) that states: in any stable queueing system,…
  • LLM as Judge — LLM-as-judge is the evaluation pattern in which one LLM scores another model's (or agent's) output against a rubric — accuracy, helpfulness, policy adherence,…
  • LLM hallucination — LLM hallucination is the failure mode where a language model "confidently makes claims that are incorrect" — it generates output that is linguistically plausible…
  • Local-remote parity — Local-remote parity is the design goal of making local development mirror the production cloud API at the API shape level, not just at the runtime-behavior level.…
  • The log is the truth, the database is a cache — The truth is the log. The database is a cache of a subset of the log. A framing — originating in Kleppmann's CIDR 2015 paper "Turning the database inside out"…
  • Logical replication — Logical replication is a replication mode in which the primary database emits a stream of row-level change events (insert / update / delete with primary-key and column values),…
  • Long-tail query — A long-tail query is a user search query that appears rarely or never in historical traffic — highly specific, uncommon, or creatively phrased.…
  • LoRA (Low-Rank Adaptation) — LoRA (Low-Rank Adaptation) is a parameter-efficient fine-tuning (PEFT) technique that freezes the weights of a pre-trained base model and trains only a small number of new…
  • LSM compaction (size-tiered, leveled, hybrid) — Log-Structured Merge (LSM) is the standard storage organisation for write-heavy systems that need to serve reasonably fast point / range / scan queries over the same data.…
  • LTAP (Lake Transactional/Analytical Processing) — LTAP (Lake Transactional/Analytical Processing) is a Databricks-coined architecture for serving both OLTP and OLAP workloads from a single governed copy of data by unifying them…
  • Machine payment — A machine payment is a payment made by an agent directly to a service, programmatically, during task execution — without a human buyer's checkout moment.…
  • Machine-readable documentation — Machine-readable documentation = repository docs structured for AI agents first, humans second: short, factual, consistently-named files with predictable layout — AGENT.md,…
  • Materialized view — A materialized view (MV) is a database object that persists the result of a query as a physical table. Unlike a regular view…
  • Matrix multiply-accumulate (MMA) — Matrix multiply-accumulate (MMA) is the fused primitive C ← A × B + C at fixed hardware-defined tile sizes, exposed by NVIDIA Tensor Cores and AMD Matrix Cores via dedicated…
  • Matryoshka Representation Learning — Matryoshka Representation Learning (MRL) is an embedding-training technique that packs information into a vector coarse-to-fine,…
  • Medallion Architecture — The Medallion Architecture is a three-tier pattern for organising data inside a data lakehouse, structured by progressive data quality. Each tier is a named "colour" of refinement:
  • Memory-bandwidth-bound — A computation is memory-bandwidth-bound when its throughput is limited by the rate at which data can be read from (or written to) main memory,…
  • Memory safety — Memory safety is the property that a program cannot access memory it isn't authorized to — no use-after-free, no buffer overrun, no double-free, no uninitialized-read.…
  • Metric cardinality — Metric cardinality is the number of unique combinations of label values a metric has — e.g., cpuusage{pod="...", tenant="..."} has one series per distinct (pod, tenant) pair.…
  • Micro-batching — Micro-batching is a stream-processing execution model that groups incoming records into small, time-bounded batches (typically 100 ms – a few seconds) and processes each batch…
  • Micro-frontends — Micro-frontends is the application of microservice-style ownership to the frontend: a page or application is composed at runtime from independently owned, independently developed,…
  • Micro-VM Isolation — Micro-VM isolation is the practice of running each tenant's code inside a minimal hardware-virtualised VM (fast boot, small memory overhead),…
  • Mixture of Experts (MoE / MMoE) — Mixture of Experts (MoE) is the neural network architecture pattern where a set of specialist subnetworks ("experts") process inputs,…
  • Model-first routing — Model-first routing is an inference-routing model where the caller specifies which model they want — a capability, e.g. "a capable reasoning model" or a named model like…
  • Model FLOPs Utilization (MFU) — Model FLOPs Utilization (MFU) is the ratio of FLOPs actually performed by a model's training or inference run to the peak theoretical FLOPs of the hardware over the same wall-clock…
  • Monolith vs microservices pendulum — The observed oscillation of engineering orgs between monolith and microservices architectures as they grow, stall, or reorganize — with neither pole being a destination.…
  • Monorepo — Monorepo = a single shared version-control repository holding many services, libraries, tools, and other artefacts that would otherwise live in separate repos.…
  • Monte Carlo simulation under uncertainty — Monte Carlo simulation evaluates a decision or policy by drawing many random samples from the underlying uncertainty distributions and averaging the outcome across samples.…
  • Multi-modal attribute extraction — Multi-modal attribute extraction is the pattern of using a vision-language model (VLM) — one that natively takes both image and text inputs…
  • Network-bound vs compute-bound — A system is network-bound when its scaling-limiting resource is network bandwidth (or packet rate) — adding more CPU / GPU does not increase throughput because bytes-per-second…
  • Network round-trip cost — The round-trip-time (RTT) floor between an application process and a remote database or RPC service is the unit cost that dominates batch-job throughput whenever a loop does one…
  • Noisy neighbor — Noisy neighbor names the multi-tenant failure mode where one tenant's workload perturbs another tenant's latency/throughput — through shared queues, shared media, shared CPU,…
  • Non-targetability — Non-targetability is the security property that an attacker cannot single out a specific individual's session, request, or storage without compromising the entire system.…
  • OAuth token lifecycle — The OAuth token lifecycle encompasses the issuance, refresh, introspection, and revocation of access and refresh tokens in an OAuth 2.0 system.
  • Object storage as disk root — A storage-architecture decision: the authoritative durability tier for a VM's disk is an S3-compatible object store, not local NVMe.…
  • Observability — The function of providing visibility into application performance and reliability via metrics, logs, and traces. The core operational quality it serves: lowering MTTD (mean time…
  • Offline-first architecture — An architectural approach that moves primary computation to the edge — designing systems to function fully without cloud connectivity and treating the cloud as a coordination,…
  • Offset-preserving replication — Offset-preserving replication is cross-cluster replication where the destination cluster holds the same per-partition offsets as the source cluster.…
  • OLTP vs OLAP — OLTP (Online Transaction Processing) and OLAP (Online Analytical Processing) are the two big workload archetypes databases are tuned for.…
  • On-Device ML Inference — On-device ML inference is the running of ML inference on end-user hardware — smartphones, laptops, browsers, embedded devices, NPUs — rather than on cloud servers.…
  • Online DDL — Online DDL is the family of techniques for applying schema changes (ALTER TABLE, CREATE INDEX, column adds, type changes,…
  • Open Table Format — An open table format (OTF) is a metadata layer over columnar data files on object storage that adds table semantics — atomic row-level updates, schema evolution,…
  • Optimistic locking — Optimistic locking is a concurrency-control discipline in which a writer reads a row (including a version column or equivalent monotonic marker),…
  • Optimistic Update — An optimistic update is a client-side UI technique where the application applies the effect of a user action to local state immediately — before the backend confirms it…
  • Out-of-band observability — Out-of-band observability is the discipline of operating a high-availability system whose internals are cryptographically or physically inaccessible to its own operators…
  • Partition skew (data skew) — Partition skew — also known as data skew — is the streaming-broker failure mode where records are distributed unevenly across a topic's partitions,…
  • Performance isolation — Performance isolation is the property that one tenant's workload does not observably affect another tenant's latency or throughput…
  • Performance prediction — Performance prediction is the problem class of estimating a system's performance metric — throughput, latency, efficiency, resource cost — from a description of its state,…
  • PID controller (feedback control) — A PID controller (proportional–integral–derivative) is a 90-year-old control-theory primitive that drives a process variable toward a setpoint by computing a correction from three…
  • Pipeline parallelism — Pipeline parallelism is a multi-GPU model-sharding strategy in which different transformer layers live on different GPUs — GPU 0 holds layers 1-N/G,…
  • Pod Disruption Budget — A Pod Disruption Budget (PDB) is a Kubernetes primitive that bounds the number or percentage of pods of a given workload that may be simultaneously terminated during a voluntary…
  • Point-in-Time Recovery (PITR) — Point-in-Time Recovery (PITR) is the database capability of producing a fresh, queryable copy of the database at a chosen past timestamp, typically for:
  • Policy-as-data — Policy-as-data stores authorization policies outside application code — in a database, a dedicated policy store, or a config bundle — so they can be changed, audited,…
  • Post-quantum authentication — Post-quantum authentication is the migration of digital- signature and credential-verification primitives to quantum- resistant alternatives…
  • Post-quantum cryptography — Post-quantum cryptography (PQC) is the class of cryptographic primitives designed to resist cryptanalytic attack by a sufficiently powerful quantum computer.…
  • Postgres logical replication slot — A Postgres logical replication slot is a row in the pgreplicationslots catalog on a Postgres primary that represents a persistent cursor into the WAL from the perspective of one…
  • Power of Two Choices (P2C) — Power of Two Choices (P2C) is a randomised load-balancing primitive that picks two backends uniformly at random for each request and then deterministically routes to the one…
  • Predicate pushdown — Predicate pushdown is the query-optimization technique of pushing filter predicates (WHERE clauses) down into the storage layer so that irrelevant data is never read into…
  • Probabilistic data structure — A probabilistic data structure trades exact correctness for sub-linear space or time by returning approximate answers with known error bounds.…
  • Progressive capability disclosure — Progressive capability disclosure is the design principle that an agent (or any consumer) should discover a system's capabilities incrementally, at the moment they are needed,…
  • Prompt injection — Prompt injection is an adversarial attack against an LLM where attacker-controlled text, embedded in input the LLM is expected to process,…
  • Q-Day — Q-Day is the day a cryptographically-relevant quantum computer (CRQC) can break the asymmetric cryptography actively protecting deployed systems — RSA, classical Diffie-Hellman,…
  • Quantization — Quantization rescales tensor elements from a high-precision floating-point range into a smaller number of discrete levels represented with fewer bits.…
  • Queueing theory (as applied to storage/IO stacks) — Queueing theory is the math of how waiting lines form and drain when arrivals are asynchronous. Applied to systems: between the CPU and durable storage there is always a chain…
  • Race condition — A race condition occurs when a system's correctness depends on the relative timing or interleaving of operations, and at least one possible ordering produces incorrect behavior.…
  • Read amplification — Read amplification is the ratio between the number of client-issued read operations and the number of substrate-level read operations they fan out into.…
  • Real user monitoring (RUM) — Real user monitoring (RUM) is a passive, inside-out observability technique: an in-page SDK captures what is actually happening in real users' browsers — Core Web Vitals,…
  • Regional failover — Regional failover is the act of shifting a service's traffic away from a failed (or degraded) geographic region into one or more healthy regions,…
  • Remote attestation — Remote attestation is the cryptographic mechanism by which a client can verify — from outside — that a specific, known-good software image is running inside a genuine TEE…
  • Remote development environment — A remote development environment is an architectural setup where the developer's editor UI runs on one machine (typically a laptop) but code execution, language services,…
  • Replicated log — A replicated log is a sequence of ordered slots, each containing a decided event, maintained identically across all functioning replicas in a distributed system.…
  • Request collapsing — Request collapsing is the CDN/cache behaviour in which N concurrent requests for the same uncached (and cacheable) resource are deduplicated into a single upstream invocation: one…
  • Retrieval-Augmented Generation (RAG) — Retrieval-Augmented Generation (RAG) is the inference-time architectural pattern where an LLM's context is augmented with documents retrieved from an external knowledge base…
  • Retrieval → ranking funnel — The retrieval → ranking funnel is the canonical two-stage architecture for recommendation, search, and recommendation-like systems at scale:
  • RPO / RTO (recovery point / time objectives) — The two canonical Disaster Recovery budget dimensions:
  • Scale to Zero — Scale-to-zero is a service-design property in which an application consumes no capacity and accrues no charge when it has no traffic,…
  • Scatter-gather query — A scatter-gather query is a query executed in a sharded system that cannot be routed to a single shard (because its predicate doesn't include the shard key or because it needs data…
  • Schema evolution — Schema evolution is the problem of changing the structure of data (records, tables, messages) over time while old and new versions coexist in the system — in flight on a queue,…
  • Schema registry — A schema registry is a centralized, versioned store of data contracts — typically event / message / record schemas — used as the single source of truth for the shape, type,…
  • Secondary index — A secondary index is a database index on a non-primary-key column (or column set). In a system with a clustered index like MySQL's InnoDB,…
  • Self-service infrastructure — The discipline of enabling engineers to provision, configure, and tear down infrastructure environments without manual ops intervention — while maintaining central governance,…
  • Separation of concerns — Separation of concerns is the practice of organising a system so that each component has a single, well-defined responsibility,…
  • Sequential vs random I/O on SSD/EBS — The same byte-level I/O demand has very different IOPS cost depending on whether the reads/writes are sequential (contiguous in the logical block address space) or random…
  • Server-Driven UI — Server-driven UI (SDUI) — sometimes backend-driven UI — is a UI architecture in which the backend decides what the client renders and what happens on interaction,…
  • Serverless Compute — Serverless compute is a model in which the provider runs customer code on demand on managed, shared infrastructure, scaling it from zero to arbitrary concurrency without…
  • Service Level Indicator (SLI) — A Service Level Indicator (SLI) is the measurement a Service Level Objective is defined over — the actual metric you collect.…
  • Shard key — A shard key is the column (or composite) whose value selects which physical shard a row lives on under horizontal sharding.…
  • Shared-nothing architecture — A shared-nothing architecture gives each node in the cluster its own private storage; no two nodes share a physical storage resource.…
  • Shared Responsibility Model — The Shared Responsibility Model is AWS's contract-level framing for which party — AWS or the customer — is responsible for what parts of a running workload.…
  • Short-lived credential auth — Short-lived credential auth is the security property that a workload's authorization is carried by credentials with minutes- scale lifetime, generated dynamically per session,…
  • Shuffle sharding — Shuffle sharding is a tenant-isolation technique that gives each tenant in a shared cluster a single-tenant experience by assigning them a randomly-chosen subset of backend nodes…
  • Side-channel attack — A side-channel attack exploits information leaked by the physical implementation of a cryptographic system rather than a weakness in the algorithm itself.…
  • Small file problem on object storage — The small file problem is the pathology of a streaming-to-lakehouse pipeline producing many small object-store files (Parquet / ORC / Avro) instead of fewer, well-sized ones.…
  • Snapshot Isolation — Snapshot Isolation (SI) is a transaction-isolation model in which each transaction reads from a consistent snapshot of the database as of its start time: it sees all writes…
  • Spatial division multiplexing — Spatial division multiplexing (SDM) is the optical-networking strategy of scaling a fiber-optic link's total capacity by increasing the number of parallel spatial paths carrying…
  • Specification-driven development — Specification-driven development is a workflow where the specification is a first-class, authored, maintained artifact — produced early, visible to customers,…
  • Speculative decoding — Speculative decoding is an LLM-inference latency-optimization technique: a small, fast drafter model proposes the next N tokens autoregressively, and a large,…
  • SSO authentication (OpenID Connect) — Single sign-on (SSO) is the pattern where a user authenticates once against an identity provider (IdP) and then uses the resulting token to prove their identity to many downstream…
  • Stale-while-revalidate cache — Stale-while-revalidate (SWR) is a caching semantic where a cache is permitted to serve a stale entry immediately while it revalidates (or re-fetches) in the background.…
  • Stateful stream processing — Stateful stream processing is a stream computation model where operators maintain mutable state across events — enabling aggregations, joins, sessionization, pattern detection,…
  • Stateless Compute — Stateless compute is a contract in which the execution environment keeps no durable state across invocations — any persistence lives in a separate managed store.…
  • Static stability — Static stability is the reliability principle that a system should continue operating with the last known good state when something fails,…
  • Step-up authentication — Step-up authentication is an access-control discipline in which a single session carries a gradient of authenticity levels,…
  • Sticky routing — Sticky routing is the property that the same logical key continues to be routed to the same node across routing-map updates, redeployments,…
  • Storage media tiering — Storage media tiering is the architectural practice of deploying multiple distinct storage media types in a coordinated hierarchy,…
  • Streaming aggregation — Streaming aggregation is the pattern of aggregating metrics in transit — as samples flow from producers to storage — instead of querying raw samples and rolling them up at read…
  • Streaming SSR — Streaming server-side rendering is an SSR variant where the server emits HTML incrementally as regions of the page become ready, instead of blocking until the full tree renders.…
  • Strong Consistency (Read-after-Write) — Strong read-after-write consistency is the guarantee that once a write to a key completes, any subsequent read of that key observes the written value…
  • Structured output reliability — Structured-output reliability is the quality axis separate from semantic correctness that asks: did the LLM produce a parseable,…
  • Submarine cable — A submarine (subsea) cable is a fiber-optic cable laid on the ocean floor to carry data between continents. It is "the least visible,…
  • Tail latency at scale — "Tail latency at scale" names the failure mode where, as a system fans a single logical operation out across N hosts, the probability that at least one host is experiencing its…
  • Tenant isolation — Tenant isolation in a multi-tenant SaaS is the property that one tenant's users, data, and policies are strictly inaccessible to other tenants…
  • Tensor parallelism — Tensor parallelism is a multi-GPU model-sharding strategy in which individual weight matrices (tensors) within each transformer layer are split across multiple GPUs…
  • Text-to-SQL — Text-to-SQL is the task of generating an executable SQL query from a natural-language question, given a specific database schema (and, in practice,…
  • Threat modeling — Threat modeling is the discipline, originating in security engineering, of enumerating threats against a system before deciding on countermeasures.…
  • Three-database problem — The three-database problem is the named infrastructure failure mode for teams building AI agents: they end up running three unrelated storage systems…
  • Thundering herd — A thundering herd is a failure mode where a resource is overwhelmed by too many simultaneous requests, typically because many clients were previously blocked / disconnected / idle…
  • Token overhead — Token overhead is the portion of an AI coding agent's per-inference cost that comes from context the user did not explicitly type — the tool outputs, codebase searches,…
  • Token vault — A token vault is an out-of-band credential-management component that holds long-lived provider credentials for the enterprise and mints short-lived,…
  • Tombstone (deletion marker) — A tombstone is a special marker written in place of a deleted record in an eventually-consistent or append-only store. It says "this key used to exist;…
  • Traffic engineering — Traffic engineering in the context of inter-domain routing refers to the deliberate manipulation of BGP attributes by an Autonomous System to influence how traffic enters and exits…
  • Training / serving boundary — The training / serving boundary is the organizational and infrastructure split between the fleet that trains a model and the fleet that serves it in production.…
  • Trusted Execution Environment (TEE) — A Trusted Execution Environment (TEE) is a hardware-enforced isolated execution context whose contents — memory, register state,…
  • Tunable consistency — Tunable consistency is the property that a database lets applications choose the consistency + durability level per operation, rather than fixing one level for the whole system.…
  • Two-tower architecture — A two-tower (or dual-encoder) model is a retrieval / ranking architecture with two independent neural encoders:
  • Valley-free routing — Valley-free routing is the property — formalised in Gao–Rexford (IEEE/ACM Trans. Networking 2001) — that a well-formed BGP AS path,…
  • Vector Embedding — A vector embedding is a dense numerical representation of a piece of unstructured data — text, image, audio, video, or document — produced by an embedding model,…
  • Vector Similarity Search — Vector similarity search is the retrieval primitive behind semantic search, recommendation, and RAG: given a query vector and a corpus of vectors (usually embeddings of documents /…
  • Vertical scaling — Vertical scaling is the discipline of increasing a single server's capacity — more CPU cores, more RAM, faster storage, higher IOPS…
  • Video transcoding — Video transcoding is the act of decoding a source video and re-encoding it to one or more target encodings, optionally changing resolution, codec, framerate, container format,…
  • VM Escape — A VM escape is when code running inside a guest virtual machine breaks through the hypervisor boundary and obtains execution or data access on the host system (or,…
  • Write-Ahead Logging (WAL) — Write-Ahead Logging (WAL) is the durability primitive under nearly every production-grade relational database, distributed log, and many object stores.…
  • Working-set memory — Working-set memory — the subset of a database's data + indexes that is actually accessed by the live query workload within a given time window and therefore needs to reside…
  • Workload identity — A workload identity is a stable, fine-grained identifier naming the logical unit of software running on a host — typically more specific than "this instance" (AMI/VM)…
  • Write amplification — Write amplification (WA) is the ratio of physical bytes written to storage to logical bytes the application intended to write.
  • Write-dependency graph — A write-dependency graph is a server-side (or client-side) data structure that tracks, for every node in an editable document,…
  • Z-Ordering — Z-Ordering is a multi-dimensional clustering technique on Delta Lake that uses a space-filling curve (the Z-order curve) to map multiple-column key tuples onto a linear ordering,…
  • Zero-knowledge proof — A zero-knowledge proof (ZKP) is a cryptographic protocol by which a prover convinces a verifier that a statement is true without revealing any information beyond the fact that it…
  • Zero-shot forecasting — Zero-shot forecasting is producing a forecast for a time series without fitting any model to that series — a pre-trained time-series foundation model takes a historical window…
  • Zero-trust authorization — Zero-trust authorization is the design principle that every tier that handles a privileged request must independently verify authorization…
Last updated · 766 distilled / 2,225 read