Improving Lakebase Postgres compute cache¶
Summary¶
Databricks details how Lakebase Postgres improves its compute-side cache — the DRAM (and local NVMe) layer that keeps hot Postgres pages close to the query engine while durable state lives in disaggregated object storage. The core change: on large fixed-size computes, retire the interim local file cache (LFC) layer and instead give Postgres shared buffers = 75% of DRAM (up from a conservative ~1 GB cap), backed by explicit 2 MB huge pages to keep address-translation overhead sane at that buffer size. This is "part 1" — fixed computes first — with a "part 2" promised for the harder autoscaling case (dynamically growing/shrinking shared buffers + huge pages in concert, contributed upstream to Postgres). The post grounds the change in three production endpoints showing ~1.3–2× throughput, up-to-5× fewer storage reads, and up to 5× lower CPU. It is a direct follow-on to the autoscaling post.
Why the standard Postgres cache is a poor fit for disaggregated storage¶
Postgres organizes rows on pages; actively accessed pages must live in the in-process shared buffers memory area. Traditionally Postgres reads pages through the OS filesystem, so the OS kernel page cache also caches those pages — a second copy. This scheme has two structural downsides and three challenges specific to a serverless, disaggregated system:
Downsides
- Double buffering. A page cached in shared buffers is also held in the OS page cache, so caching 1 GB of data can consume ~2 GB of RAM — halving effective cache capacity on the compute. (Source: sources/2026-09-10-databricks-improving-lakebase-postgres-compute-cache)
- Kernel eviction is blind to Postgres. The OS page cache knows nothing about shared buffers or DB internals, so it cannot make smart page- replacement decisions.
Technical challenges
- In a disaggregated storage system, reads from storage do not travel through the OS filesystem / page cache at all — so the "free" kernel cache layer that classic Postgres leans on simply isn't in the path.
shared_buffersis a static (boot-time) parameter — it can't change without restarting Postgres, which is hostile to a serverless autoscaling system.- Process-per-connection. Postgres forks an OS process per active connection; each maps shared buffers into its own address space, so larger shared buffers multiply per-connection memory-management (page table) cost.
The Lakebase cache path — LFC as the interim answer¶
To work around the static shared_buffers limit and the missing kernel
cache, Lakebase introduced a local file cache (LFC): an autoscaling
secondary cache on the compute node's local NVMe, presented to users as
one high-speed compute cache but internally up to two tiers:
- Shared buffers — Postgres's in-memory buffer; lowest-latency path.
- Local file cache (LFC) — larger-capacity cache on local NVMe; higher capacity than DRAM but requires disk I/O to serve a page.
A miss across both tiers is routed from the compute node to the
distributed storage layer (a
GetPage request). Shared buffers were tuned conservatively — capped at
1 GB so a compute running at minimum CU wouldn't over-consume memory —
with LFC filling out the rest (up to 75% of DRAM total). The downside: on
large working sets the 1 GB shared-buffers cap forced most cache hits down
into the slower NVMe LFC tier instead of DRAM. The stated intent is to
retire LFC's current form as Lakebase progresses toward fully dynamic
shared buffers.
Larger shared buffers (fixed computes, live today)¶
First delivery targets fixed-size computes (because shared buffers aren't yet dynamic): disable the LFC and set shared buffers to 75% of DRAM. Live today for fixed-size computes with CU ≥ 80. Eliminating the ~1 GB cap keeps hot pages in the fastest DRAM layer instead of cascading to NVMe. This also kills double buffering (Lakebase reads bypass the OS page cache, so 1 GB cached now costs 1 GB, not 2 GB) and lets eviction decisions be made with knowledge of DB state — the door to smarter-than-OS replacement policies.
- Verify via
show shared_buffersin a Postgres connection; an 80 CU endpoint should report15278640(≈ 15.3 M × 8 KB blocks ≈ 116 GiB).
Huge pages — taming address-translation overhead at large buffer sizes¶
Sizing shared buffers to 75% of DRAM ran headfirst into challenge #3 (process-per-connection). Each backend maps shared buffers into its own address space and needs its own page table entries (PTEs) — the kernel structures the CPU walks to translate virtual → physical addresses. Linux defaults to 4 KB pages:
- Each 1 GB of shared buffers ⇒ 262,144 PTEs per process.
- At 32 GB shared buffers × 512 backends ⇒ ~4.3 billion PTEs ≈ 32 GB of page tables to map 32 GB of cache.
That working set dwarfs the CPU's Translation Lookaside Buffer (TLB), so even a shared-buffers hit pays a penalty from TLB misses + page-table walks. The Postgres community's remedy is huge pages (2 MB each): switching to huge pages cuts page-table size by 512× and sharply lowers TLB miss rates. Databricks' benchmarks: huge pages reduced tail read latency by up to ~40% and CPU utilization by up to ~30%.
Huge pages in virtualized environments¶
Lakebase Postgres runs in lightweight guest VMs on bare-metal hosts, so address translation crosses two virtualized layers; huge pages only pay off if implemented consistently host reservation → hypervisor VM backing → guest kernel (a break at any tier degrades the benefit). Databricks added dedicated huge-page backing across VM infra, choosing explicit 2 MB HugeTLB pages over best-effort transparent huge pages (THP). VMs for large fixed-size computes initialize with a predetermined volume of huge pages sufficient for Postgres startup; compute startup releases surplus huge pages beyond what Postgres needs.
- Verify via
show huge_pages; an 80 CU endpoint should report"on".
Part 2 — autoscaling shared buffers (roadmap)¶
Bringing large shared buffers to autoscaling computes is harder: shared buffers must grow on scale-up and shrink on scale-down while allocating the exact huge-page volume needed at each size. Databricks built a protocol to autoscale huge pages provided to the guest, scaled in concert with dynamic shared buffers to preserve efficient translation at high concurrency/memory. Part 2 will cover the dynamic-shared-buffers implementation, the state of open-source Postgres here, and the areas Databricks is advancing and contributing upstream.
Key takeaways¶
- Disaggregated storage removes the kernel page cache from the read path, so the classic shared-buffers + OS-page-cache scheme both wastes RAM (double buffering) and loses its second cache tier — a serverless Postgres has to own its DRAM caching explicitly. (Source: sources/2026-09-10-databricks-improving-lakebase-postgres-compute-cache)
- A conservative shared-buffers cap is a hidden tax on large working sets. Lakebase's 1 GB cap pushed most hits into the slower NVMe LFC tier; the fix is to give Postgres up to 75% of DRAM as shared buffers and drop the interim LFC on large fixed computes.
- Big shared buffers are impossible without huge pages under Postgres's process-per-connection model: 32 GB buffers × 512 backends ≈ 4.3 B 4 KB PTEs ≈ 32 GB of page tables. 2 MB HugeTLB pages cut that 512× → ~40% lower tail read latency, ~30% lower CPU in benchmarks.
- Huge pages must be consistent through every virtualization tier (host → hypervisor → guest); Databricks chose explicit HugeTLB over THP and pre-provisions huge pages at VM init, releasing surplus at startup.
- Production results (region-by-region rollout, Aug 2026): Example 1 —
~2× throughput, storage
GetPage/sdropped ~8K → ~1.5K (≈5× fewer storage reads), lower p50/p99. Example 2 — ~1.3× (43%) throughput, compute cache hit rate ≈ 100% (served almost entirely from shared buffers). Example 3 — CPU fell 20 cores → 4 (5× lower), hit rate ≈ 100%, 2× throughput. - Incremental delivery on a serverless roadmap. Fixed computes shipped first (static shared buffers make them tractable); autoscaling (dynamic shared buffers + dynamically scaled huge pages) is part 2, with fixes contributed upstream to Postgres.
Operational numbers¶
- Shared buffers: previously capped at 1 GB; now 75% of DRAM on fixed
computes with CU ≥ 80 (LFC disabled). 80 CU →
shared_buffers = 15278640(8 KB blocks ≈ 116 GiB),huge_pages = on. - Huge page size: 2 MB (explicit HugeTLB, not THP). 4 KB → 2 MB reduces page-table size 512×.
- PTE math: 1 GB buffers ⇒ 262,144 PTEs/process; 32 GB × 512 backends ⇒ ~4.3 B PTEs ≈ 32 GB of page tables.
- Benchmark deltas from huge pages: tail read latency −~40%, CPU −~30%.
- Production (large endpoints): ~2× and ~1.3× and ~2× throughput across three endpoints; storage reads ~8K→~1.5K GetPage/s; CPU 20→4 cores; compute cache hit rate ≈ 100%.
Caveats¶
- Large shared buffers are live only for fixed-size computes with CU ≥ 80; autoscaling computes still use the LFC path pending part 2.
- The
15278640/"on"verification values are for an 80 CU endpoint specifically; other sizes differ. - Benchmark figures (~40% tail-latency, ~30% CPU) are Databricks' own internal benchmarks; production examples are three selected large endpoints, not a fleet-wide distribution.
- LFC is being retired in its current form, not deleted outright — it still backs computes that haven't moved to large shared buffers.
Source¶
- Original: https://www.databricks.com/blog/improving-lakebase-postgres-compute-cache
- Raw markdown:
raw/databricks/2026-09-10-improving-lakebase-postgres-compute-cache-2939f6d2.md
Related¶
- System: systems/lakebase — the serverless Postgres whose compute cache this improves
- System: systems/postgresql — shared buffers, process-per-connection, huge pages
- System: systems/pageserver-safekeeper — the disaggregated storage layer misses route to
- Concept: concepts/compute-storage-separation — the forcing function for explicit DRAM caching
- Concept: concepts/working-set-memory — what shared buffers must hold to stay in DRAM
- Concept: concepts/cache-hit-rate — the ≈100% hit-rate outcome
- Prior post: sources/2026-08-31-databricks-autoscaling-lakebase-postgres — the autoscaling context this extends