Autoscaling Lakebase Postgres¶
Summary¶
Databricks describes how Lakebase Postgres eliminates instance sizing entirely through autoscaling — resizing a running Postgres compute in both directions, on a live database, without dropping connections. The architectural precondition is storage/compute separation: because durable state (WAL + pages) lives in the storage layer and the compute VM owns no durable state, a compute node can start, stop, move, or change size without moving the database underneath it. The post breaks autoscaling into two problems — when to resize (the algorithm) and how to resize a running VM without stopping Postgres (the mechanism). The algorithm tracks three independent signals (CPU, memory, and compute-cache working set), each producing its own target size, and takes the largest. The most novel piece is a time-windowed modification of HyperLogLog that estimates the current working set (not all-time cardinality) so the compute cache can be sized to the data actively in use. The mechanism uses NeonVM (a QEMU/KVM Kubernetes controller) to add/remove CPU and memory in place, with live migration as the fallback when a node is too full.
Key takeaways¶
-
Storage/compute separation is the enabling foundation, not a detail. Traditional Postgres ties execution and durable state to one machine, so resizing is a database operation. Lakebase's compute layer "owns no durable state" — RAM + local NVMe only — while the storage layer (safekeepers on SSD replicate WAL; pageservers reconstruct page versions; object storage holds the long-term immutable record) owns durability and history. A compute node can therefore "start, stop, move, or change size without moving the database underneath it." (Source: this article; storage deep-dive is a companion post on object storage + WAL.)
-
Three autoscaling signals, largest wins — each produces its own target compute size (in Compute Units, CU), and the final target is
max(cpuGoalCU, memGoalCU, lfcGoalCU), then clamped to the user-configured min/max autoscaling limits. See three-signal-largest-wins-autoscaling. -
CPU signal (
cpuGoalCU) — every 5 seconds the autoscaler-agent reads the VM's 1-minute CPU load average; the goal keeps load at or below 90% of available CPU. The 1-minute average filters short spikes; the 5-second poll tracks the moving average. CPU alone is insufficient: a query stalled on a network page fetch shows low CPU while performing poorly. -
Memory signal (
memGoalCU) runs at two frequencies because memory has a harsher failure mode than CPU (OOM-kill vs slowdown): the autoscaler-agent reads VM-wide memory every 5 seconds, and an in-VM vm-monitor checks Postgres memory every 100 milliseconds. Goal keeps usage below 75% of allocated RAM. The vm-monitor also vetoes any downscale that would leave running processes without enough memory. History: this polling design replaced an earlier cgroupmemory.high-event design — crossingmemory.highmade Linux reclaim memory and throttle the cgroup; polling proved more predictable and stable while still giving a 100ms view. -
Compute-cache signal (
lfcGoalCU) measures whether the workload's active data fits close to Postgres. The compute cache (originally the Local File Cache / LFC) is a disk-backed cache sized to fit in the kernel page cache, acting as a resizable extension of Postgres shared buffers; when compute grows the vm-monitor expands the cache into the added memory. This closes the CPU-only blind spot: cache misses leave queries waiting on network, which reduces CPU, so a CPU-only autoscaler would shrink exactly when a bigger cache would help. -
Time-windowed HyperLogLog is the working-set estimator. A standard HyperLogLog only grows and can only answer "how many distinct pages since Postgres started" — useless for autoscaling, which needs "how many distinct pages belong to the workload running now." Lakebase's modification: instead of setting a bit when a page-hash is observed, store the current timestamp at that register. To estimate cardinality since time
T, treat registers updated afterTas set and older ones as unset. This produces a distinct-page estimate for any window ending now (last 1 min, 5 min, 60 min). Every 20 seconds the autoscaler-agent collects working-set estimates for windows from 1 to 60 minutes. -
Choosing the window is itself adaptive. No universal window describes a database's current working set: too short discards cache aggressively between bursts; too long keeps memory allocated for finished work. The algorithm watches how the estimate changes as the window extends — a steady workload plateaus (more time, few new pages), while a recently-ended heavy workload shows a jump once the window reaches back far enough to include it. The search for that jump starts after 5 minutes (preventing shrink during a short pause); if no sharp increase is found it uses the 60-minute estimate (the expected result for a stable hour-long workload).
-
Cache-growth projection avoids "measuring too late." Measuring the current working set lands slightly late — if the cache grows only after new pages are read, early pages may already have been evicted, forcing a re-fetch. So the algorithm projects working-set growth forward, examining how the estimate increases duration-to-duration and allocating enough cache for the working set expected by the next control interval. Because cache metrics are fetched every 20s the projection covers only a fraction of a minute — longer projections would react earlier but amplify spikes and cause oscillation. The projected size becomes
lfcGoalCU, goaled to fit the working set within up to 75% of the compute's RAM. See working-set-projection-for-cache-sizing. -
Four components coordinate a resize: (a) autoscaler-agent (per Kubernetes node) collects metrics, computes targets, initiates scaling; (b) vm-monitor (in each VM) watches Postgres memory, validates downscales, resizes the compute cache; (c) a modified Kubernetes scheduler — the single source of truth for allocation — must approve every upscale before memory is committed; (d) NeonVM (a custom Kubernetes resource + controller built on QEMU/KVM) applies CPU/memory changes to a running guest.
-
Scale-up sequence: agent computes target → scheduler checks the node can satisfy it without overcommitting memory → agent updates the NeonVM resource → NeonVM controller adds CPU/memory to the running VM → vm-monitor expands the compute cache. The scheduler seeing both ordinary Kubernetes scheduling and autoscaling requests is essential; otherwise it could place a new workload on a node at the same moment the autoscaler committed the remaining memory to a Postgres VM.
-
Live migration is the escape hatch when a node is full. If a node is too full to grow in place, NeonVM can live-migrate the VM to another node; the VM keeps its IP so existing connections stay open. Because Lakebase computes hold little durable local state, migration is mostly VM memory + runtime state.
-
Scaling down is treated as importantly as scaling up. "Some autoscaling systems are quick to add capacity but slow to give it back, leaving databases oversized long after a spike has passed. Lakebase Postgres treats both directions the same way." Downscale reuses the same components with one extra in-VM check (the vm-monitor confirms enough memory remains for Postgres
- guest before proceeding). See symmetric-up-down-autoscaling.
Operational numbers¶
- Control loop runs at three timescales: 100 ms (vm-monitor checks Postgres memory for rapid allocation), 5 s (autoscaler-agent reads CPU + overall memory), 20 s (autoscaler-agent evaluates working-set estimates across 1–60 min windows).
- CPU goal target: ≤ 90% of available CPU (1-minute load average).
- Memory goal target: < 75% of allocated RAM (VM-wide).
- Compute-cache goal: fit the working set within ≤ 75% of the compute's RAM.
- Working-set window search starts at 5 min, falls back to the 60-min estimate if no jump is detected.
- A production database can change size more than 32,000 times per month.
- Final scaling target:
max(cpuGoalCU, memGoalCU, lfcGoalCU)clamped to user-configured min/max limits.
Caveats / undisclosed¶
- The exact CU→(vCPU, RAM) mapping and step granularity of in-place resize is not specified.
- The register/precision parameters of the modified HyperLogLog (number of
registers
m, resulting relative error) are not disclosed. - How compute-cache coherence is maintained across a live migration (does the cache survive the move, or re-warm on the destination?) is not addressed.
- The scheduler's overcommit / bin-packing policy beyond "don't overcommit memory" is not detailed.
- NeonVM lineage note (verbatim from the post): "Lakebase Postgres architecture started in Neon and the resource/controller name remains the same." So NeonVM is a Neon-origin component now used inside Lakebase.
Source¶
- Original: https://www.databricks.com/blog/autoscaling-lakebase-postgres
- Raw markdown:
raw/databricks/2026-08-31-autoscaling-lakebase-postgres-b59b1be7.md - Companion storage deep-dive: Object storage + WAL in the agentic era
Related¶
- systems/lakebase — the host system; this is its autoscaling layer.
- systems/neon — architectural origin (NeonVM, pageserver/safekeeper).
- systems/neonvm — the QEMU/KVM Kubernetes controller that applies resizes.
- systems/pageserver-safekeeper — the durable storage tier that makes stateless-compute resize possible.
- in-place-vm-resize — resize a running guest without restart.
- concepts/probabilistic-data-structure — the working-set estimator.
- concepts/working-set-memory — the quantity being estimated and cached.
- three-signal-largest-wins-autoscaling — the target-selection rule.
- working-set-projection-for-cache-sizing — projecting cache growth.
- symmetric-up-down-autoscaling — scale down as fast as up.