SYSTEM Cited by 1 source
lakebase_text¶
lakebase_text is a Postgres extension from Databricks that brings native BM25 full-text search to Lakebase Postgres. Generally available on AWS and Azure (2026-09-28), it is the full-text counterpart to lakebase_vector; combining the two yields native hybrid search inside Postgres. (Source: sources/2026-09-28-databricks-lakebase-search)
The problem it replaces: tsvector lacks corpus context¶
Standard Postgres full-text search (tsvector + a GIN index) lacks
corpus-wide relevance context — it matches terms but cannot weight them by how
rare or informative they are across the whole document collection. lakebase_text
scores terms with global inverse document frequency (IDF): heavily weighting
rare, high-intent terms while penalizing common filler words — the defining
property of BM25 over naive term matching. (Source:
sources/2026-09-28-databricks-lakebase-search)
Why it is faster than tsvector + GIN¶
lakebase_text is faster than traditional tsvector + GIN indexes because it
evaluates score upper-bounds during traversal to skip entire posting blocks
that cannot reach the top-K results — a block-max / WAND-style dynamic-pruning
optimization common in modern IR engines (Lucene's BlockMaxWAND), where the
engine maintains per-block maximum scores and short-circuits blocks that provably
cannot enter the result set. This turns a full posting-list scan into a pruned
traversal. (Source: sources/2026-09-28-databricks-lakebase-search)
Hybrid search: fuse with lakebase_vector + SQL¶
Combining lakebase_text with lakebase_vector unlocks native hybrid search inside Postgres. In a single query you can:
- Fuse semantic vector search with BM25 keyword relevance (patterns/parallel-retrieval-fusion);
- Apply standard SQL filter predicates;
- Join directly against live operational tables.
This replaces a multi-system retrieval pipeline (separate search engine + ETL) with one governed SQL call, keeping lexical exact-match (acronyms, IDs, error strings) and semantic paraphrase recall in the same database.
Seen in¶
- sources/2026-09-28-databricks-lakebase-search — launch/architecture post; canonical source for the global-IDF BM25 design, the tsvector+GIN critique, the block-skipping (upper-bound pruning) speedup, and native hybrid search.
Related¶
- systems/lakebase — host Postgres.
- systems/lakebase-vector — the ANN vector sibling.
- systems/bm25 — the scoring algorithm implemented.
- concepts/hybrid-search — the combined vector + lexical retrieval mode.
- concepts/retrieval-ranking-funnel — where BM25 sits in a retrieval stack.