The 40-year-old database rule agents just broke: how LTAP unifies OLTP and OLAP workloads¶
Summary¶
A Databricks interview with Jonathan Katz (longtime Postgres contributor, Senior Staff PM at Databricks) framing LTAP (Lake Transactional/Analytical Processing) as the answer to a 40-year separation between operational (OLTP) and analytical (OLAP) database systems. Where the earlier Matei Zaharia post described LTAP's storage mechanics bottom-up, this interview argues the motivation: autonomous AI agents (e.g. fraud detection) act on data settling in hundreds of milliseconds and can't tolerate the batch-copy delay of the classic pipeline-and-copy model, while a fleet of agents running analytical scans can overwhelm an OLTP engine that was never built for them. LTAP unifies the two worlds at the storage layer, not the engine — Lakebase's PageServer flushes durable Postgres data to object storage and represents it in the same columnar Parquet/Iceberg/Delta format the Lakehouse already uses, so Spark/SQL read live operational data directly with no CDC pipeline and no load on the transactional system. All of it sits under one Unity Catalog governance boundary.
Key takeaways¶
- Two worlds, rooted in physics, not preference. Row storage returns a single fast answer (OLTP); columnar storage scans/aggregates across everything (OLAP). Analytical queries are designed to consume the whole parallel system for one answer over a large dataset; operational queries are designed to consume as few resources as possible while still answering fast. These opposite optimization targets are why the two systems have always sat apart. (Source: this article)
- Agents are the workload that finally breaks the separation. A fraud/anomaly agent works in hundreds of milliseconds against a system also handling a huge volume of writes and short reads; a herd of agents can overwhelm an operational system without guardrails. (Source: this article)
- LTAP = analytics on live operational data, no movement, no load. "Lake Transactional/Analytical Processing lets you run analytical queries directly against live operational data without moving that data anywhere and without putting load on the system serving your transactions." (Source: this article)
- Only possible because Lakebase decouples compute from storage. Inherited from Neon: stateless, ephemeral Postgres compute; a durable storage layer that absorbs high-throughput writes and periodically flushes to object storage independent of whatever compute is running. (concepts/compute-storage-separation, concepts/stateless-compute)
- Bit-preserving row→columnar transcoding is the hard part. Postgres has its own data types and encodings; Iceberg/Delta have theirs. Databricks writes the data into Parquet "without changing a single bit" of the original Postgres physical representation, merging the operational and analytical representations of the same data into one copy. (Source: this article)
- Two-tier storage: hot rows, cool columns. The storage layer runs a hotter tier in row format for fast operational access and a cooler tier in columnar format for analytical reads, so either side gets its preferred layout. (concepts/storage-media-tiering — hot/cool access-frequency axis)
- Why LTAP succeeds where HTAP stalled: separate serverless compute per workload. HTAP works but is "expensive… clunky, hard to run, and generally not open." LTAP gives serverless operational compute and serverless analytical compute as two independently-scaled things. "Storage is the cheap part of any data system. Compute is the expensive part" — so unify the cheap layer and keep a specialized engine for each job. (Source: this article)
- The stale-data failure mode, concretely. Fraud detection on a batch copy minutes/hours old is too slow to catch a transaction clearing in milliseconds. Point the agent at the operational system instead and a full-purchase-history anomaly scan is an expensive OLAP query hitting an OLTP engine — degrading every other transaction trying to clear (concepts/noisy-neighbor); a fleet of such agents overwhelms it fast. (Source: this article)
- Governance is the other half: one catalog over both worlds. The Lakehouse innovation was a centralized, unified view — who can access what, consistent policies (e.g. SSN visible only to a privileged group). Operational systems were historically data silos where, once you shipped data downstream via a pipeline, "nobody owned what happened to it." LTAP puts all data under one catalog so operational data never leaves a governance boundary just to be analyzed. (concepts/governed-agent-data-access)
- Open foundation as a design principle. Built on Postgres (closing on the #3 most-favored database on DB-Engines). With LTAP's unified storage you stop moving data to get the right engine on it — "you're bringing the engine to the data."
- Collapses the bronze/silver pipeline tax. "I don't need to run pipelines anymore just to get data into a bronze or silver layer. I can start analyzing it the moment it's written." (concepts/medallion-architecture)
- The one-sentence framing. The core shift is unified storage: bring the operational and analytical representations of the same data together "without moving anything, without any bits changing." LTAP doesn't erase the difference between a transaction and an analytical query — it removes the tax of running both against the same data.
Systems / concepts / patterns extracted¶
- Systems: systems/lakebase, systems/postgresql, systems/neon, systems/pageserver-safekeeper, systems/apache-parquet, systems/apache-iceberg, systems/delta-lake, systems/unity-catalog, systems/apache-spark.
- Concepts: concepts/ltap (canonical name for the storage-layer unification idea), concepts/oltp-vs-olap, concepts/compute-storage-separation, concepts/columnar-storage-format, concepts/storage-media-tiering (hot/cool tiers), concepts/stateless-compute, concepts/medallion-architecture, concepts/noisy-neighbor, concepts/ai-agent-guardrails, concepts/governed-agent-data-access, concepts/open-table-format, concepts/change-data-capture (the approach LTAP replaces).
- Patterns: patterns/tiered-storage-to-object-store (durable writes flushed to object storage), patterns/presentation-layer-over-storage (one storage copy, two engine "views" — row for OLTP, columnar for OLAP).
Operational numbers / claims¶
- Fraud-relevant transactions clear in hundreds of milliseconds or less — the latency budget an agent must beat, which batch copies cannot.
- Two storage tiers: hot (row) + cool (columnar) over the same durable data.
- Zero bits changed in Postgres→Parquet transcoding (lossless physical representation preserved).
- Postgres is closing on #3 on DB-Engines' most-favored-database ranking (cited as a signal, explicitly not an adoption measure).
Caveats¶
- Interview / thought-leadership format: qualitative argument, few hard numbers (quantitative claims — 5×/2×/>10× — live in the companion Zaharia post, sources/2026-06-30-databricks-from-monolith-to-lakebase-to-ltap).
- Vendor post: LTAP is a Databricks/Lakebase capability, framed favorably vs HTAP; "rolling out" maturity caveats from the companion post still apply.
- The bit-preserving row→columnar transcoding is novel; production maturity at scale is asserted, not independently benchmarked here.
Source¶
- Original: https://www.databricks.com/blog/40-year-old-database-rule-agents-just-broke-how-ltap-unifies-oltp-and-olap-workloads
- Raw markdown:
raw/databricks/2026-09-07-the-40-year-old-database-rule-agents-just-broke-how-ltap-uni-74cbd35d.md
Related¶
- concepts/ltap — the canonical concept this article co-establishes
- sources/2026-06-30-databricks-from-monolith-to-lakebase-to-ltap — companion post, storage mechanics bottom-up
- systems/lakebase — the serverless Postgres LTAP is built on
- systems/neon — storage/compute-separated Postgres Lakebase descends from
- systems/unity-catalog — the governance boundary over both worlds
- concepts/oltp-vs-olap — the 40-year separation LTAP bridges
- concepts/compute-storage-separation — the enabling architectural property
- concepts/columnar-storage-format — the analytical layout, Parquet
- concepts/storage-media-tiering — hot (row) / cool (columnar) tiers
- concepts/change-data-capture — the pipeline LTAP eliminates