Skip to content

CONCEPT Cited by 13 sources

Compute–storage separation

Compute–storage separation is the architectural property where a system's persistence layer and its query/compute layer are decoupled and scale independently. Storage is a durable, shared substrate (object store, shared log, distributed KV); compute is a pool of workers that can be added/removed without touching the data.

Why it matters

  • Scale the expensive resource independently. Aggregation workloads are compute-heavy but bounded by data size; you can burst compute up for the job without over-provisioning storage, and vice-versa.
  • Multiple compute tiers against one source of truth. Serving, analytics, backfill, and disaster-recovery workers can all read the same storage without replica sprawl.
  • Elastic pricing. Compute goes to zero when idle; storage bill stays flat. Aligns with a pay-per-use model (see elasticity).
  • Horizontal aggregation. An OLAP engine with separated compute can parallelise an aggregation across many workers that all read the same storage; most of the calculation ends up in memory.

Contrast with shared-nothing OLTP

Classic MySQL / Postgres setups bundle compute and storage on the same box. To scale, you grow the box (Canva doubled RDS instance size every 8–10 months until the model broke) or shard the data, which complicates the application. OLAP warehouses like Snowflake split these layers: storage lives in the cloud object layer, compute is a "virtual warehouse" cluster you size independently. Canva explicitly calls this out as the reason Snowflake can aggregate billions of records in a few minutes, "several orders of magnitude faster" than the MySQL round-trip approach. (Source: sources/2024-04-29-canva-scaling-to-count-billions)

Seen elsewhere

  • S3 + analytics engines — S3 is the durable store; engines like Spark, Trino, Athena, and Snowflake run compute over the same object data. (systems/aws-s3, systems/apache-iceberg)
  • Aurora DSQL — separates the journal (durability) from the adjudicator/crossbar/storage/execution tiers, each scaling on its own axis. (systems/aurora-dsql)
  • Lakebase — Neon-descended serverless Postgres; externalises page and WAL storage into systems/pageserver-safekeeper, leaving the Postgres compute VMs ephemeral and scale-to-zero. A cleaner example of the OLTP-shape separation than DSQL (which keeps Postgres's page/WAL layer and rewrites replication/concurrency). (systems/lakebase)
  • Lambda — stateless compute with state pushed to managed stores is the same separation at the request-level. (concepts/stateless-compute)

Caveats

  • Query latency ≠ OLTP latency. Compute is elastic but not free to spin up; cold queries can be seconds. OLAP stores aren't a serving tier — see warehouse-unload-bridge.
  • Network between layers matters. Separation implies moving data or metadata between compute and storage; engine designs work hard on caching, local-disk staging, and predicate push-down to keep that in check.
  • Consistency story gets more complex. The storage layer must give the compute layer enough consistency guarantees for correctness; S3 moving to strong read-after-write consistency in 2020 is one enabling step for this (see concepts/strong-consistency).

Seen in

  • sources/2026-10-01-databricks-lakebase-postgres-branch-based-restores-for-fast-recovery-at-scale — Separation makes restore a metadata operation instead of a data-movement job. Databricks 2026-10-01. The traditional coupled-monolith restore is slow because compute and storage ship together: you provision a new instance, hydrate the snapshot from object storage onto the Postgres volume, then replay WAL — each stage scales with data size. In Lakebase, because compute runs Postgres while the Safekeeper + Pageserver + object-storage tier owns durability and history, "a restore does not copy data into a new disk … it simply creates a branch at a timestamp which is a simple metadata operation, not a multi-hour copy-and- replay job." The payoff attributable directly to the split: restore time is independent of data size ("100 TB as fast as 10 GB"). Reinforces the recurring theme that separating durability from compute reshapes a whole class of operations (branch, restore, scale-to-zero, HA) from copies into pointer/ metadata work.

  • sources/2026-09-10-databricks-improving-lakebase-postgres-compute-cache — separation reshapes caching, not just durability: because Lakebase reads bypass the OS filesystem entirely, the kernel page cache that classic Postgres relies on is not in the path, so the compute must own its DRAM caching explicitly (shared buffers = 75% of DRAM, backed by huge pages) instead of leaning on the OS. A concrete example that "compute owns no durable state" also means "compute owns its own cache hierarchy" (DRAM → local NVMe → object storage), with the storage layer as the authoritative record and compute memory as pure cache.

  • sources/2026-09-07-databricks-the-40-year-old-database-rule-agents-just-broke-how-ltap-unifies-oltp-and-olap-workloads — Katz: compute–storage separation (inherited from Neon) is the precondition for LTAP. Because durable state is decoupled from ephemeral Postgres compute, the storage layer can flush writes to object storage and transcode them to columnar Parquet, then serve serverless operational compute and serverless analytical compute as two independently-scaled things — "storage is the cheap part… compute is the expensive part."

  • sources/2026-05-07-databricks-how-lakebase-architecture-delivers-5x-faster-postgres-writes — First canonical wiki instance of compute-storage separation enabling structural elimination (not just relocation) of a durability primitive. Classical Postgres's Full Page Write exists to tolerate torn pages on the local-disk page heap. On Lakebase's stateless-compute-streaming-WAL-to- safekeeper-quorum architecture, "there is no local-disk page to tear, the failure mode FPW was designed to prevent simply does not exist" — FPW is structurally redundant on compute. Databricks disabled compute-side FPW across the global fleet; the incidental read-path reset-point role FPW had on the pageserver's delta-chain replay is preserved by image-generation-pushdown-to-storage — image generation moved from the compute's WAL stream to the storage tier's background processing, with per-page-threshold cadence replacing checkpoint-scoped FPW cadence. First wiki disclosure of the Paxos-based safekeeper quorum as the durability substrate that makes stateless compute structurally safe. Quantified outcomes: 94% compute WAL-volume reduction; 5× write throughput at 32 vCPU; p99 read latency −30% to −50%. This ingest extends the four prior Lakebase sources into a five-axis canonicalisation of compute-storage separation at serverless-Postgres altitude: (a) forcing function for two-tier encryption + CMK revocation (2026-04-20); (b) operational payoff for bursty workloads — no data movement on compute spin-up (2026-04-27); (c) enabler for agent-driven per-request lifecycle (2026-04-29); (d) enabler for PITR as a routine sub-10-second operation (2026-04-30); (e) enabler for elimination of durability primitives specific to local-disk failure modes (2026-05-07).

  • sources/2026-04-30-databricks-backstage-with-lakebase — Fourth Lakebase-altitude confirmation; new axis: PITR-as-storage-fork-cheap-enough-to-be-routine. Thoughtworks Backstage POC demonstrates that concepts/point-in-time-recovery collapses into a copy-on-write storage fork at a past timestamp only because the storage tier is independent of compute + keeps historical page versions as durable shared state. Measured: 3.78 seconds end-to-end PITR for a 32-row deletion recovery; same primitive as forward-branching (1.09 s for 63 MB Backstage catalog), differing only in the source_branch_time parameter. Canonicalises the architectural unification [[patterns/branching-is-pitr- with-time-now]]: on a compute-storage-separated substrate with copy-on-write storage, branching and PITR are the same operation with different time inputs. Extends the four Lakebase sources into a four-axis canonicalisation of compute-storage separation at serverless-Postgres altitude: (a) forcing function for two-tier encryption

  • CMK revocation (2026-04-20); (b) operational payoff for bursty workloads — no data movement on compute spin-up (2026-04-27); (c) enabler for agent-driven per-request lifecycle (2026-04-29); (d) enabler for PITR as a routine sub-10-second operation rather than last-resort disaster recovery (2026-04-30). The common thread: decoupling storage from compute doesn't just scale-independently — it makes every storage-addressing operation (clone, fork, recover, historical-query) structurally cheap because compute is stateless + disposable.

  • sources/2026-04-29-databricks-and-stripe-projects-infrastructure-built-for-agents — Compute-storage separation as agent-lifecycle enabler. Third canonical cross-source confirmation on the wiki that compute-storage separation is the load-bearing property behind Lakebase's serverless- Postgres contract. Verbatim: "By decoupling compute from storage, agents can create, build, and tear down OLTP databases in seconds." New axis the prior confirmations didn't cover: the separation as the forcing function for per-request compute lifecycle at agent-initiated cadence — not just bursty workload tolerance (LangGuard 2026-04-27) or two-tier encryption (CMK 2026-04-20). Creating an OLTP database becomes a sub-350 ms operation because storage already exists and only compute needs standing up; tearing one down is cheap because only compute is de-allocated. This makes database instances structurally disposable — the pre-condition for agent-provisioned-database as a resource primitive. The three Lakebase sources now form a coherent canonicalisation of compute-storage separation at serverless-Postgres altitude: (a) forcing function for two-tier encryption + CMK revocation semantics (2026-04-20); (b) operational payoff for bursty workloads — no data movement on compute spin-up (2026-04-27); (c) enabler for agent-driven per-request lifecycle (2026-04-29).

  • sources/2026-04-27-databricks-inside-one-of-the-first-production-deployments-of-lakebase-langguard — LangGuard on Lakebase: the second production-deployment datapoint for Lakebase's compute/storage-separated Postgres. Canonicalises the specific operational payoff for bursty agentic workloads: "Because durable state lives in the storage layer, not in the compute node, spinning up a new compute instance requires no data movement. It simply attaches to the existing database history and begins serving queries immediately." No cold-start penalty on the data side — only on the compute-VM side — which is what makes scale-to-zero between bursts operationally viable for a latency-sensitive-enforcement workload. Former IBM QRadar team frames this as the "answer we had always needed but didn't have access to" when building SIEM at petabyte scale.

  • sources/2024-04-29-canva-scaling-to-count-billions — Snowflake's separated compute is the reason Canva's end-to-end recompute is feasible at billions-per-month scale.
  • sources/2026-04-20-databricks-take-control-customer-managed-keys-for-lakebase-postgres — Lakebase's Pageserver (durable pages) + Safekeeper (durable WAL) storage tier is independent of the ephemeral Postgres compute VMs that scale up/down/to-zero; the separation is explicitly called out as the forcing function motivating a two-tier encryption design (persistent envelope + per-boot ephemeral keys).
  • sources/2024-07-29-aws-amazons-exabyte-scale-migration-from-apache-spark-to-ray-on-ec2 — Amazon Retail's post-Oracle BI stack named as the canonical application of compute-storage separation: S3 for storage, swappable compute engines on top (systems/amazon-redshift, systems/aws-rds, systems/apache-hive, systems/apache-spark on systems/amazon-emr, later systems/ray on systems/aws-ec2). The architectural payoff is named explicitly: swapping compute engines is architecturally cheap because the source of truth — S3 — doesn't move. BDT's migration of the compactor workload from Spark to Ray for the largest ~1% of tables is the case study — storage untouched, compute engine replaced, 82% better cost per GiB.
  • — PlanetScale's structural pushback: compute-storage separation via EBS imposes a variance floor that translates directly into customer-visible partial failure on OLTP workloads ("this is the cost of separating storage and compute and the sheer complexity of the software and networking components between the client and the backing disks"). The post argues OLTP databases should go the other way — re-collocate compute and storage on direct-attached NVMe + pay for durability via cluster replication. See systems/planetscale-metal + shared-nothing-storage-topology + direct-attached-nvme-with-replication. Canonical wiki datapoint against compute-storage separation as the default for low-latency OLTP.

  • — Brian Morrison II (PlanetScale, 2024-01-24). Canonical OLTP illustration of compute-storage separation via Amazon Aurora: the writer compute node and read-only compute nodes share the same distributed storage substrate (10 GiB segments × 6 copies × 3 AZs, 4-of-6 write-ack quorum). "Since data is replicated on the storage level, read-only compute nodes can be started at any time in an availability zone containing a copy of the data for that node to read." Unlike the analytics/warehouse instances (Snowflake, Databricks), Aurora runs the separation at the OLTP tier with writer-driven page-update notifications to keep reader buffer caches coherent. See storage-forwarded-redo-log-replication for the full substrate pattern. This is the structural counter- model to PlanetScale's shared-nothing direct-attached-NVMe substrate — both architectures ship under the Aurora-style "primary + read replicas" surface but use opposite substrate primitives underneath.

  • sources/2026-09-28-databricks-lakebase-search — Compute-storage separation as the enabler for object-storage-resident vector search. lakebase_vector puts the ANN index (IVF cluster blocks + RaBitQ 1-bit codes) in cheap object storage, with RAM/NVMe as ephemeral caches — the same Lakebase substrate that makes compute disposable now makes the vector index disposable and scale-to-zero-able (P90 1.13 s cold start; 100M vectors on 1 CU). This is the structural answer to pgvector/HNSW, whose RAM-resident index has no working-set notion. Extends the Lakebase compute-storage-separation thread from durability/branching/LTAP into the search-index axis.

  • sources/2026-10-01-cloudflare-announcing-cloudflare-k2-serverless-event-streams — separation applied to a streaming broker. Cloudflare K2 stores its entire log in R2, so compute (the edge services that accept produces and serve consumes) and storage (R2) scale independently. Cloudflare names the benefit explicitly: "A secondary benefit is that it separates compute and storage, meaning each can be scaled independently. This allows us to store vast quantities of historical data at low cost." The sharper claim than usual: offloading replication and consensus (not just bytes) to the storage layer is what lets the compute tier stay "radically simpler, cheaper, and higher performance" — the opposite of Kafka's stateful brokers that own both compute and replicated storage.

  • sources/2026-10-01-cloudflare-introducing-cloudflare-basin-an-open-serverless-data-platform — separation as the openness thesis of an analytics platform. Basin (Cloudflare's GA data platform) elevates compute-storage separation from an implementation property to the product pitch: "separating the storage layer from the compute layer and allowing you to use the right query engine for the job." Storage is Iceberg on R2; compute is any Iceberg-compatible engine (PyIceberg, DuckDB, Snowflake, Spark, or Cloudflare's own Basin SQL). The novel point: free egress is what makes the separation economically real — data portability across tools/regions/clouds "is only possible with free egress" (concepts/egress-cost). Basin SQL is the on-platform OLAP compute tier that scatter-gathers queries across Workers against the separated R2 storage.

  • concepts/oltp-vs-olap

  • concepts/elt-vs-etl
  • concepts/elasticity
  • concepts/stateless-compute
  • systems/snowflake, systems/aws-s3, systems/aurora-dsql

  • sources/2026-09-30-databricks-a-practical-guide-to-cost-optimization-with-lakebase-postgre — Separation of storage and compute framed explicitly as the root cost lever. Databricks' Lakebase cost guide makes the economic argument concrete: because branches, read replicas, and high availability all add compute that reads the same underlying storage layer, "scaling read capacity does not require creating and paying for another copy of the database," and HA "adds redundant compute across availability zones while continuing to use the existing highly available storage layer." The same property makes copy-on-write branching cheap (pay for divergence, not a full copy) and makes scale-to-zero safe (persistent state stays in the storage tier, so suspending compute does not lose the database). Canonical statement that storage/compute separation turns normally-expensive duplication (replicas, HA, dev copies) into compute-only adds.

Last updated · 766 distilled / 2,225 read