Skip to content

DATABRICKS 2026-08-31

Read original ↗

Autoscaling Lakebase Postgres

Summary

Databricks describes how Lakebase Postgres eliminates instance sizing entirely through autoscaling — resizing a running Postgres compute in both directions, on a live database, without dropping connections. The architectural precondition is storage/compute separation: because durable state (WAL + pages) lives in the storage layer and the compute VM owns no durable state, a compute node can start, stop, move, or change size without moving the database underneath it. The post breaks autoscaling into two problems — when to resize (the algorithm) and how to resize a running VM without stopping Postgres (the mechanism). The algorithm tracks three independent signals (CPU, memory, and compute-cache working set), each producing its own target size, and takes the largest. The most novel piece is a time-windowed modification of HyperLogLog that estimates the current working set (not all-time cardinality) so the compute cache can be sized to the data actively in use. The mechanism uses NeonVM (a QEMU/KVM Kubernetes controller) to add/remove CPU and memory in place, with live migration as the fallback when a node is too full.

Key takeaways

  1. Storage/compute separation is the enabling foundation, not a detail. Traditional Postgres ties execution and durable state to one machine, so resizing is a database operation. Lakebase's compute layer "owns no durable state" — RAM + local NVMe only — while the storage layer (safekeepers on SSD replicate WAL; pageservers reconstruct page versions; object storage holds the long-term immutable record) owns durability and history. A compute node can therefore "start, stop, move, or change size without moving the database underneath it." (Source: this article; storage deep-dive is a companion post on object storage + WAL.)

  2. Three autoscaling signals, largest wins — each produces its own target compute size (in Compute Units, CU), and the final target is max(cpuGoalCU, memGoalCU, lfcGoalCU), then clamped to the user-configured min/max autoscaling limits. See three-signal-largest-wins-autoscaling.

  3. CPU signal (cpuGoalCU) — every 5 seconds the autoscaler-agent reads the VM's 1-minute CPU load average; the goal keeps load at or below 90% of available CPU. The 1-minute average filters short spikes; the 5-second poll tracks the moving average. CPU alone is insufficient: a query stalled on a network page fetch shows low CPU while performing poorly.

  4. Memory signal (memGoalCU) runs at two frequencies because memory has a harsher failure mode than CPU (OOM-kill vs slowdown): the autoscaler-agent reads VM-wide memory every 5 seconds, and an in-VM vm-monitor checks Postgres memory every 100 milliseconds. Goal keeps usage below 75% of allocated RAM. The vm-monitor also vetoes any downscale that would leave running processes without enough memory. History: this polling design replaced an earlier cgroup memory.high-event design — crossing memory.high made Linux reclaim memory and throttle the cgroup; polling proved more predictable and stable while still giving a 100ms view.

  5. Compute-cache signal (lfcGoalCU) measures whether the workload's active data fits close to Postgres. The compute cache (originally the Local File Cache / LFC) is a disk-backed cache sized to fit in the kernel page cache, acting as a resizable extension of Postgres shared buffers; when compute grows the vm-monitor expands the cache into the added memory. This closes the CPU-only blind spot: cache misses leave queries waiting on network, which reduces CPU, so a CPU-only autoscaler would shrink exactly when a bigger cache would help.

  6. Time-windowed HyperLogLog is the working-set estimator. A standard HyperLogLog only grows and can only answer "how many distinct pages since Postgres started" — useless for autoscaling, which needs "how many distinct pages belong to the workload running now." Lakebase's modification: instead of setting a bit when a page-hash is observed, store the current timestamp at that register. To estimate cardinality since time T, treat registers updated after T as set and older ones as unset. This produces a distinct-page estimate for any window ending now (last 1 min, 5 min, 60 min). Every 20 seconds the autoscaler-agent collects working-set estimates for windows from 1 to 60 minutes.

  7. Choosing the window is itself adaptive. No universal window describes a database's current working set: too short discards cache aggressively between bursts; too long keeps memory allocated for finished work. The algorithm watches how the estimate changes as the window extends — a steady workload plateaus (more time, few new pages), while a recently-ended heavy workload shows a jump once the window reaches back far enough to include it. The search for that jump starts after 5 minutes (preventing shrink during a short pause); if no sharp increase is found it uses the 60-minute estimate (the expected result for a stable hour-long workload).

  8. Cache-growth projection avoids "measuring too late." Measuring the current working set lands slightly late — if the cache grows only after new pages are read, early pages may already have been evicted, forcing a re-fetch. So the algorithm projects working-set growth forward, examining how the estimate increases duration-to-duration and allocating enough cache for the working set expected by the next control interval. Because cache metrics are fetched every 20s the projection covers only a fraction of a minute — longer projections would react earlier but amplify spikes and cause oscillation. The projected size becomes lfcGoalCU, goaled to fit the working set within up to 75% of the compute's RAM. See working-set-projection-for-cache-sizing.

  9. Four components coordinate a resize: (a) autoscaler-agent (per Kubernetes node) collects metrics, computes targets, initiates scaling; (b) vm-monitor (in each VM) watches Postgres memory, validates downscales, resizes the compute cache; (c) a modified Kubernetes scheduler — the single source of truth for allocation — must approve every upscale before memory is committed; (d) NeonVM (a custom Kubernetes resource + controller built on QEMU/KVM) applies CPU/memory changes to a running guest.

  10. Scale-up sequence: agent computes target → scheduler checks the node can satisfy it without overcommitting memory → agent updates the NeonVM resource → NeonVM controller adds CPU/memory to the running VM → vm-monitor expands the compute cache. The scheduler seeing both ordinary Kubernetes scheduling and autoscaling requests is essential; otherwise it could place a new workload on a node at the same moment the autoscaler committed the remaining memory to a Postgres VM.

  11. Live migration is the escape hatch when a node is full. If a node is too full to grow in place, NeonVM can live-migrate the VM to another node; the VM keeps its IP so existing connections stay open. Because Lakebase computes hold little durable local state, migration is mostly VM memory + runtime state.

  12. Scaling down is treated as importantly as scaling up. "Some autoscaling systems are quick to add capacity but slow to give it back, leaving databases oversized long after a spike has passed. Lakebase Postgres treats both directions the same way." Downscale reuses the same components with one extra in-VM check (the vm-monitor confirms enough memory remains for Postgres

    • guest before proceeding). See symmetric-up-down-autoscaling.

Operational numbers

  • Control loop runs at three timescales: 100 ms (vm-monitor checks Postgres memory for rapid allocation), 5 s (autoscaler-agent reads CPU + overall memory), 20 s (autoscaler-agent evaluates working-set estimates across 1–60 min windows).
  • CPU goal target: ≤ 90% of available CPU (1-minute load average).
  • Memory goal target: < 75% of allocated RAM (VM-wide).
  • Compute-cache goal: fit the working set within ≤ 75% of the compute's RAM.
  • Working-set window search starts at 5 min, falls back to the 60-min estimate if no jump is detected.
  • A production database can change size more than 32,000 times per month.
  • Final scaling target: max(cpuGoalCU, memGoalCU, lfcGoalCU) clamped to user-configured min/max limits.

Caveats / undisclosed

  • The exact CU→(vCPU, RAM) mapping and step granularity of in-place resize is not specified.
  • The register/precision parameters of the modified HyperLogLog (number of registers m, resulting relative error) are not disclosed.
  • How compute-cache coherence is maintained across a live migration (does the cache survive the move, or re-warm on the destination?) is not addressed.
  • The scheduler's overcommit / bin-packing policy beyond "don't overcommit memory" is not detailed.
  • NeonVM lineage note (verbatim from the post): "Lakebase Postgres architecture started in Neon and the resource/controller name remains the same." So NeonVM is a Neon-origin component now used inside Lakebase.

Source

  • systems/lakebase — the host system; this is its autoscaling layer.
  • systems/neon — architectural origin (NeonVM, pageserver/safekeeper).
  • systems/neonvm — the QEMU/KVM Kubernetes controller that applies resizes.
  • systems/pageserver-safekeeper — the durable storage tier that makes stateless-compute resize possible.
  • in-place-vm-resize — resize a running guest without restart.
  • concepts/probabilistic-data-structure — the working-set estimator.
  • concepts/working-set-memory — the quantity being estimated and cached.
  • three-signal-largest-wins-autoscaling — the target-selection rule.
  • working-set-projection-for-cache-sizing — projecting cache growth.
  • symmetric-up-down-autoscaling — scale down as fast as up.
Last updated · 766 distilled / 2,225 read