Skip to content

DATABRICKS 2026-09-28

Read original ↗

Lakebase Search: State-of-the-art full text and vector search for Postgres

Summary

Databricks brings a scalable search engine to Lakebase Postgres via two GA extensions: lakebase_vector (approximate nearest-neighbor / ANN vector search) and lakebase_text (BM25 full-text search), both generally available on AWS and Azure. The core argument is that pgvector — the most-installed extension in Lakebase — hits a wall at scale because HNSW indexes must live entirely in RAM to be fast, making cost scale with data volume rather than usage, making index builds slow and blocking, and preventing per-query parallelism. lakebase_vector rebuilds the vector index around Lakebase's storage/compute separation: durable data rests in cheap object storage while RAM/NVMe act as ephemeral caches, and the index is built from hierarchical IVF clustering plus binary RaBitQ quantization (~1 bit/dim, ~32× smaller than float32) so it is fast both hot-in-RAM and cold-on-object-storage. On the VectorDBBench 100M LAION benchmark it claims 2× the throughput of the next best system, 4× cheaper than a cloud Postgres vendor on pgvector, P99 latency of 71 ms at 97% recall, and scale-to-zero with a P90 first-query-after-idle of 1.13 s. lakebase_text adds corpus-aware BM25 (global IDF) that beats tsvector + GIN, and combining the two gives native hybrid search fused with SQL filter predicates and joins in a single query. (Source: this article.)

Key takeaways

  1. pgvector's structural wall is RAM-residency, not algorithm quality. Because HNSW search is random-access graph traversal, queries are fast only when the whole index fits in RAM; the moment it spills to disk, queries become chains of random reads and performance drops 10×–50×. A 768-dim float32 vector needs ~3.3 KB after graph links + Postgres overhead, so 100M rows ≈ 330 GB RAM to stay resident — with no working-set notion: you provision for the entire index whether you query all of it or none of it. (Source: this article.) See concepts/working-set-memory.

  2. Index maintenance blocks the database. A pgvector HNSW build can take ~50 hours on a standard cloud instance when it spills to disk (millions of random I/Os). Inserts are slow because each write navigates and modifies multiple graph layers via random-access lookups, and — because HNSW lacks global rebalancing — restoring quality requires a full REINDEX that locks the table and blocks production writes. (Source: this article.)

  3. pgvector can't parallelize a single query. Each pgvector query runs on one Postgres backend process, so the HNSW scan is never parallelized across cores. Higher recall means visiting more graph nodes → more random reads → higher latency and lower QPS; the only throughput lever is more connections or read-replicas. (Source: this article.)

  4. lakebase_vector = hierarchical IVF clustering + binary RaBitQ quantization over object storage. Vectors are grouped into IVF clusters stored as contiguous blocks; a query scores centroids in memory, then reads only the few promising blocks — turning hundreds of random hops into a handful of large sequential reads (concepts/sequential-vs-random-io). Each vector is compressed with RaBitQ binary quantization to ~1 bit/dimension (~32× smaller); queries scan the compact codes to shortlist candidates, then rerank the bounded shortlist against full-precision vectors (patterns/cheap-approximator-with-expensive-fallback). (Source: this article.)

  5. Statelessness enables true scale-to-zero for search. Decoupling storage from compute makes lakebase_vector stateless: a node caches hot data on demand, suspends to zero when idle, resumes on the next query. At rest you pay only for storage; cold starts are cheap because only the quantized codes and the specific blocks a query touches hydrate — measured P90 1.13 s for the first query after scale-to-zero (100M × 768-dim). It can serve 100M vectors on just 1 Lakebase Compute Unit (CU). (Source: this article.) See concepts/scale-to-zero, concepts/cold-start.

  6. Index builds are parallel and offloadable. Centroids are trained once on a small random sample (the only whole-dataset step); after that every vector is independently assigned to its nearest centroid, quantized, and written into its cluster block — fanning across as many cores as available. Databricks' LTAP architecture (open formats) lets it offload index builds off the primary database entirely to distributed engines like Spark, bringing build times down to minutes ("stay tuned"). (Source: this article.)

  7. Filtering is inline during the cluster-block scan. lakebase_vector applies SQL predicates as it scans cluster blocks, avoiding over-fetching candidates and keeping recall high on filtered queries — a predicate-pushdown-style optimization for ANN. (Source: this article.)

  8. lakebase_text brings corpus-aware BM25 to Postgres, beating tsvector+GIN. Standard Postgres tsvector lacks corpus-wide relevance context; lakebase_text scores terms with global inverse document frequency (IDF) — weighting rare, high-intent terms and penalizing common filler. It is faster than tsvector + GIN because it evaluates score upper-bounds during traversal to skip entire posting blocks that can't reach top-K (a WAND / block-max-style optimization). (Source: this article.) See systems/bm25.

  9. Native hybrid search in one SQL query. Combining lakebase_text with lakebase_vector lets a single query fuse semantic vector search with BM25 keyword relevance, apply SQL filter predicates, and join directly against live operational tables — replacing a multi-system retrieval pipeline with one governed SQL call (concepts/hybrid-search, patterns/parallel-retrieval-fusion). (Source: this article.)

  10. Designed for agentic burstiness. Agents made vector search a core stack requirement and introduced extreme burstiness — a single workflow can trigger thousands of concurrent retrievals in seconds. Lakebase Search targets this with a serverless architecture that scales "from 1 row to 1 billion vectors and from 1 QPS to thousands" without manual reprovisioning. (Source: this article.)

Operational numbers

Metric Value Notes
VectorDBBench 100M throughput 2× next best system LAION 100M dataset
Cost vs cloud pgvector vendor 4× cheaper before autoscaling savings
Recall / latency P99 71 ms @ 97% recall 100M LAION
Vector size (768-dim float32) ~3.3 KB after graph links + PG overhead
RAM for 100M pgvector index ~330 GB to keep resident for ms queries
pgvector disk-spill slowdown 10×–50× random-read chains
pgvector HNSW build time ~50 hours standard cloud instance, disk-spill
RaBitQ compression ~1 bit/dim, ~32× vs float32
Scale-to-zero cold start P90 1.13 s first query, 100M × 768-dim
Density 100M vectors on 1 CU Lakebase Compute Unit
Customer (Conexiom) 3× lower DB spend, 5× throughput BM25 hybrid over 100M+ rows, ½ compute footprint vs pgvector

Caveats & unknowns

  • Vendor marketing post. Benchmark numbers (2× throughput, 4× cheaper, P99 71 ms) are Databricks' own; VectorDBBench is a public harness but the comparison configuration is partly self-selected. For pgvector and DiskANN, the post notes they "only tested performance on a single large instance" — the comparison is not apples-to-apples on a distributed footprint. See concepts/benchmark-methodology-bias.
  • Spark-offloaded index builds are "stay tuned" — the minutes-scale distributed build is described as roadmap, not shipped GA.
  • No disclosure of the exact IVF fan-out (nprobe), number of clusters, RaBitQ reranking shortlist size, or how filtered recall degrades vs unfiltered.
  • Databricks positions Databricks AI Search (managed, out-of-the-box retrieval, see systems/mosaic-ai-vector-search) as the better fit when you want quality without tuning; Lakebase Search is for consolidating operational + search data in one database.

Extracted entities

Source

Last updated · 766 distilled / 2,225 read