Skip to content

CONCEPT Cited by 2 sources

LTAP (Lake Transactional/Analytical Processing)

LTAP (Lake Transactional/Analytical Processing) is a Databricks-coined architecture for serving both OLTP and OLAP workloads from a single governed copy of data by unifying them at the storage layer instead of the engine. A transactional Postgres engine and an analytical Lakehouse engine read the same durable data — the operational side sees it as rows, the analytical side sees it as columns — with no CDC pipeline, no data movement, and no analytical load on the transactional system.

The name deliberately rhymes with HTAP, but the design choice is the opposite: HTAP tries to build one engine good at both jobs; LTAP keeps a specialized, mature engine for each and shares only the (cheap) storage substrate.

The problem it targets

For ~40 years, operational (OLTP) and analytical (OLAP) systems have been separate because of storage physics: row layout answers a single query fast, columnar layout scans and aggregates fast, and the two optimization targets are irreconcilable in one physical representation (concepts/oltp-vs-olap, concepts/columnar-storage-format). Bridging them meant a pipeline: copy operational data into a warehouse. That copy is stale, it requires per-table opt-in, and "nobody owned what happened to it" downstream of a governance boundary. (Source: sources/2026-09-07-databricks-the-40-year-old-database-rule-agents-just-broke-how-ltap-unifies-oltp-and-olap-workloads)

AI agents are framed as the workload that finally forces the issue: a fraud/anomaly agent must decide in the hundreds of milliseconds a card transaction takes to clear, so a batch copy minutes/hours old is useless — but pointing the agent at the operational system means running an expensive full-history scan against an engine tuned for the opposite, degrading every concurrent transaction (concepts/noisy-neighbor). A fleet of such agents overwhelms the operational system fast without guardrails. (Source: sources/2026-09-07-databricks-the-40-year-old-database-rule-agents-just-broke-how-ltap-unifies-oltp-and-olap-workloads)

How it works (Lakebase implementation)

LTAP is only possible on top of compute–storage separation. In systems/lakebase (descended from Neon):

  • Postgres compute is stateless and ephemeral; durable state lives in the SafeKeeper/PageServer tier on object storage (concepts/stateless-compute).
  • As the PageServer materializes pages to object storage, spare CPU transcodes row data into columnar Parquet in an open table format (Delta/Iceberg). The Postgres physical representation is preserved "without changing a single bit"; MVCC row versions are retained for Postgres but invisible to Iceberg/Delta readers.
  • Analytical engines (Spark, Photon) read bulk data straight from object storage. For freshness they ask Postgres only for the current LSN (a cheap metadata op), then fetch just the unmaterialized delta from the PageServer — so the OLTP system serves near-zero analytical I/O. (Source: sources/2026-06-30-databricks-from-monolith-to-lakebase-to-ltap)

Two-tier storage

The storage layer runs a hotter tier in row format for fast operational access and a cooler tier in columnar format for analytical reads — the access-frequency axis of concepts/storage-media-tiering applied to a single logical dataset, so each side gets its preferred layout. (Source: sources/2026-09-07-databricks-the-40-year-old-database-rule-agents-just-broke-how-ltap-unifies-oltp-and-olap-workloads)

Why LTAP over HTAP

  • Cost of compute vs storage. "Storage is the cheap part of any data system. Compute is the expensive part." Unify the cheap layer; keep a specialized engine per workload. LTAP gives serverless operational compute and serverless analytical compute as two independently-scaled things, instead of one expensive engine trying to be good at both.
  • Performance isolation by construction. Workloads are isolated at the compute level while sharing storage — the isolation HTAP historically failed to provide.
  • Mature ecosystems. Reuses full-featured Postgres + Lakehouse engines rather than a new engine with an incomplete feature set / no ecosystem (the classic HTAP failure modes).
  • Governance travels with the data. All tables sit under one Unity Catalog boundary, so operational data never leaves governance just to be analyzed (concepts/governed-agent-data-access).
  • Every table analytical by construction. Unlike CDC/mirroring (per-table opt-in, replication lag), LTAP makes all tables queryable analytically with no pipeline to build or monitor — collapsing the bronze/silver ingest tax of concepts/medallion-architecture (concepts/change-data-capture).

The slogan: instead of moving data to the right engine, "bring the engine to the data."

Contrast

  • vs HTAP — same goal (one system, both workloads), opposite mechanism: engine-unification (HTAP) vs storage-unification (LTAP).
  • vs CDC to a warehouse — LTAP removes the copy and the lag; the analytical view is the operational data, one LSN behind at most.
  • vs concepts/oltp-vs-olap — LTAP does not erase the row/column archetype split; it keeps both engines and only removes the tax of running them against the same data.

Caveats

  • Vendor-specific: LTAP is a Databricks/Lakebase capability, announced "rolling out in coming months," maturity-at-scale asserted rather than independently benchmarked.
  • The bit-preserving row→columnar transcoding is novel; very small tables are not converted to columnar form (an optimization trade-off).

Seen in

Merged aliases

  • storage-layer-unification-over-engine-unification
  • storage-layer-unification
Last updated · 766 distilled / 2,225 read