SYSTEM Cited by 1 source
VectorDBBench¶
VectorDBBench is an open-source benchmark harness for vector databases and ANN indexes, measuring the throughput (QPS), latency, and recall tradeoff of a system on standard vector datasets. In this wiki it appears as the benchmark Databricks uses to position lakebase_vector against pgvector and DiskANN. (Source: sources/2026-09-28-databricks-lakebase-search)
Why it exists¶
Vector-search systems trade recall against cost and latency, and every vendor can pick an operating point that flatters its system. A shared harness on a common dataset — here the LAION 100M dataset (100 million image/text embeddings) — makes the recall-vs-QPS-vs-cost curve comparable across systems, which is why it shows up as the reference in vendor benchmarking claims. See concepts/vector-similarity-search for the underlying recall/latency tradeoff.
Databricks' reported results (LAION 100M)¶
lakebase_vector on VectorDBBench LAION 100M (Source: sources/2026-09-28-databricks-lakebase-search):
- 2× the throughput of the next best system.
- 4× cheaper than a cloud Postgres vendor using pgvector (before autoscaling savings).
- P99 latency 71 ms at 97% recall.
Caveat — read the configuration¶
Benchmark numbers here are the vendor's own, and the post notes that for pgvector and DiskANN it only tested performance on a single large instance — so the comparison is not necessarily apples-to-apples across a distributed footprint. Treat as directional vendor evidence, not neutral third-party measurement. See concepts/benchmark-methodology-bias for the general hazard of vendor-run benchmarks (dataset choice, operating-point selection, hardware parity).
Seen in¶
- sources/2026-09-28-databricks-lakebase-search — LAION 100M results cited to position lakebase_vector against pgvector and DiskANN.
Related¶
- systems/lakebase-vector — the system under test.
- concepts/vector-similarity-search — recall/latency/QPS tradeoff being measured.
- concepts/benchmark-methodology-bias — why vendor benchmark claims need scrutiny.