Skip to content

CONCEPT Cited by 6 sources

Hybrid Search

Definition

Hybrid search combines lexical retrieval (BM25 / keyword / exact-term) with semantic retrieval (dense vector similarity) — usually running both in parallel and fusing their result lists — so that a single query gets both the exact-term precision of keyword search and the paraphrase/synonym recall of vector search. It is the dominant production retrieval shape for RAG and agentic search, precisely because neither mode alone is sufficient:

  • Lexical (BM25) nails acronyms, proper nouns, IDs, and specific error strings — cases where embedding similarity drifts.
  • Semantic (vectors) covers paraphrase ("bought a bicycle" ≈ "purchased a bike") — cases BM25's exact-token matching misses.

See concepts/retrieval-ranking-funnel for how hybrid retrieval sits inside a broader retrieve → fuse → rerank funnel.

Fusion methods

The two ranked lists are combined by one of:

  • Reciprocal Rank Fusion (RRF) — score each doc by Σ 1/(k + rank_i) across lists; robust, no score calibration needed.
  • Weighted score sum — normalize and blend the raw scores (requires calibrating lexical vs vector score scales).
  • Learned reranker — a cross-encoder or gradient-boosted model re-scores the union of candidates. See patterns/parallel-retrieval-fusion.

Why it is a first-class production requirement

Multiple production systems keep BM25 as a primary surface rather than a legacy fallback, and layer vectors alongside it:

Native hybrid search inside the database

Historically hybrid search meant stitching a dedicated search engine to the primary database with an ETL pipeline. Databricks' Lakebase Search collapses this into Postgres itself: combining lakebase_text (BM25) with lakebase_vector (ANN) lets a single SQL query fuse semantic + keyword relevance, apply SQL filter predicates, and join directly against live operational tables — "one SQL tool call for agents" replacing a multi-system retrieval pipeline. (Source: sources/2026-09-28-databricks-lakebase-search) This is the in-database realization of hybrid search: the fusion happens where the operational data already lives, governed by standard database rules.

Seen in

Last updated · 766 distilled / 2,225 read