SYSTEM Cited by 1 source
RaBitQ¶
RaBitQ is a binary quantization algorithm for vectors that compresses each vector to roughly 1 bit per dimension (~32× smaller than float32) while preserving enough fidelity to shortlist nearest-neighbor candidates, which are then reranked at full precision. Introduced in the paper "RaBitQ: Quantizing High-Dimensional Vectors with a Theoretical Error Bound for Approximate Nearest Neighbor Search" (arXiv 2405.12497). It is the quantization scheme underpinning Databricks' lakebase_vector. (Source: sources/2026-09-28-databricks-lakebase-search)
What it does¶
RaBitQ is a specific point in the quantization design space — the extreme low-bit end (binary / 1-bit) applied to vector-search codes rather than model weights. Its distinguishing property versus naive 1-bit sign quantization is a theoretical error bound on the estimated distance, which makes it usable as a reliable shortlisting step: the compact codes give an approximate distance good enough to select a bounded candidate set, and the true ranking is recovered by reranking that shortlist against the full-precision vectors.
In lakebase_vector this composes with IVF clustering:
- Vectors are grouped into IVF cluster blocks; centroids are scored in memory.
- Within promising blocks, the ~1-bit RaBitQ codes are scanned to broaden the candidate set cheaply.
- The bounded shortlist is reranked against full-precision vectors — cheap approximator with an expensive, bounded fallback.
The compression is what makes the index small enough to be cheap both hot in RAM and cold on object storage: "When cached, search operates over a tiny footprint using quantized vectors." (Source: sources/2026-09-28-databricks-lakebase-search)
Why binary quantization matters for object-storage vector search¶
Keeping full-precision float32 vectors resident is exactly the cost driver that sinks RAM-bound indexes like HNSW (a 768-dim float32 vector is ~3 KB). At ~1 bit/dimension a 768-dim vector's code is ~96 bytes, so the entire scannable code set for 100M vectors is small enough to hydrate cheaply on demand — the enabling factor behind lakebase_vector's scale-to-zero cold-start (P90 1.13 s) and its ability to serve 100M vectors on 1 compute unit. Full-precision vectors are only touched for the final rerank of a tiny shortlist.
Seen in¶
- sources/2026-09-28-databricks-lakebase-search — named as the binary quantization used by lakebase_vector (~1 bit/dim, ~32× smaller than float32), with the scan-codes-then-rerank flow.
Related¶
- concepts/quantization — the general technique; RaBitQ is the binary / 1-bit, vector-search-code variant.
- systems/lakebase-vector — consumer.
- concepts/ivf-inverted-file-index — the clustering it composes with.
- patterns/cheap-approximator-with-expensive-fallback — the shortlist + rerank shape.