Skip to content

SYSTEM Cited by 1 source

lakebase_vector

lakebase_vector is a Postgres extension from Databricks that provides scalable approximate nearest-neighbor (ANN) vector search inside Lakebase Postgres. Generally available on AWS and Azure (2026-09-28), it is positioned as the successor to pgvector for large-scale vector workloads — the same SQL surface, but an index architecture built around Lakebase's storage/compute separation rather than requiring the whole index to fit in RAM. (Source: sources/2026-09-28-databricks-lakebase-search)

The problem it replaces: pgvector's RAM wall

pgvector is the most-installed extension in Lakebase Postgres and provides ANN via HNSW and IVFFlat over native Postgres data. Databricks names three structural pain points at scale (Source: sources/2026-09-28-databricks-lakebase-search):

  1. Cost scales with data volume, not usage. HNSW is fast only when the whole index is RAM-resident (random-access graph traversal). Spill to disk → chains of random reads → 10×–50× slowdown. A 768-dim float32 vector needs ~3.3 KB after graph links + overhead, so 100M rows ≈ 330 GB RAM. There is no working-set notion (concepts/working-set-memory) — you provision for the entire index regardless of query pattern.
  2. Index maintenance is expensive and blocks the DB. A disk-spilling HNSW build can take ~50 hours; inserts are slow (each write modifies multiple graph layers via random access); and HNSW lacks global rebalancing, so restoring quality requires a full REINDEX that locks the table and blocks production writes.
  3. No per-query parallelism. Each pgvector query runs on a single Postgres backend, so the HNSW scan never parallelizes across cores. Higher recall = more graph nodes visited = higher latency / lower QPS.

Architecture — IVF clusters + RaBitQ over object storage

Lakebase separates storage from compute: durable data rests in cheap cloud object storage, while RAM and local NVMe are ephemeral caches in front of it (concepts/object-storage-as-disk-root). An HNSW cache in this model would be a series of random object-store reads. lakebase_vector instead uses an index that is fast both hot-in-RAM and cold-on-object-storage, via two ideas:

  • Hierarchical IVF clustering. Vectors are grouped into clusters stored as contiguous blocks. A query scores the cluster centroids in memory, then reads only the few promising blocks — turning hundreds of random hops into a handful of large sequential reads (concepts/sequential-vs-random-io).
  • Binary quantization via RaBitQ (arXiv 2405.12497). Each vector is compressed to ~1 bit per dimension (~32× smaller than float32). Queries scan the compact codes to shortlist candidates, then rerank that bounded shortlist against the full-precision vectors — a cheap-approximator + expensive-rerank funnel.

When cached, search operates over a tiny footprint using quantized vectors. When cold, queries fetch only the blocks they need — no crawl of the entire index.

Properties delivered

  • Pay-per-use, scale-to-zero. Storage/compute decoupling makes the extension stateless: a node caches hot data on demand, suspends to zero when idle, and resumes on the next query (concepts/scale-to-zero). At rest you pay only for storage. Cold starts are cheap because only quantized codes + the specific touched blocks hydrate — measured P90 1.13 s for the first query after scale-to-zero (100M × 768-dim, concepts/cold-start). It can serve 100M vectors on just 1 Lakebase Compute Unit (CU).
  • Fast, offloaded index builds. Centroids are trained once on a small random sample (the only whole-dataset step); afterward each vector is independently assigned to its nearest centroid, quantized, and written to its cluster block — fanning across all available cores. Because data sits in open formats, Lakebase's LTAP architecture can offload index builds off the primary database to distributed engines like Spark, bringing build times down to minutes (roadmap — "stay tuned").
  • Fast, accurate, parallel search. Broadens candidate search cheaply using 1-bit codes, reranks only a tight shortlist at full precision. Because index blocks are independent, a single query parallelizes across CPU cores — breaking pgvector's single-backend limitation — delivering high recall and low latency simultaneously.
  • Inline filtering. SQL predicates are applied as lakebase_vector scans cluster blocks, avoiding over-fetching candidates and keeping recall high on filtered queries (concepts/predicate-pushdown).

Benchmark claims

On the VectorDBBench LAION 100M dataset (Source: sources/2026-09-28-databricks-lakebase-search):

  • 2× the throughput of the next best system.
  • 4× cheaper than a cloud Postgres vendor using pgvector (before autoscaling savings).
  • P99 latency 71 ms at 97% recall.

Customer Conexiom reports running BM25 hybrid search over 100M+ rows with half the compute footprint of their prior pgvector setup, and 3× lower DB spend at 5× throughput. (For pgvector and DiskANN the post tested only a single large instance — see concepts/benchmark-methodology-bias.)

Relationship to other vector indexes

  • Versus HNSW (pgvector default): HNSW is RAM-bound graph ANN; lakebase_vector is IVF + quantization over object storage, so it does not require the whole index in RAM and parallelizes a single query.
  • Versus DiskANN: both target larger-than-RAM, but DiskANN is a pure graph on SSD; lakebase_vector uses IVF cluster blocks + binary quantization with in-memory centroid scoring.
  • Versus SPFresh/SPANN (PlanetScale): a parallel industry answer to larger-than-RAM transactional vector indexes — SPANN also uses partitioned posting lists on SSD; lakebase_vector's distinctive bet is object-storage residency + scale-to-zero + Spark-offloaded builds.
  • Versus Databricks AI Search: the managed, tuning-free retrieval product; lakebase_vector is the in-database option when you want operational + search data unified in Postgres.

Seen in

Last updated · 766 distilled / 2,225 read