Open-sourcing Metals v2: Databricks' Java and Scala language server for multi-million line codebases¶
Summary¶
Databricks forked and reworked Metals, the widely-used
open-source Scala language server, to add first-class Java support and to scale
low-latency code intelligence across a 26M-line Bazel monorepo (24M lines
Scala, plus Java and Protobuf across 142k+ files). The motivating context is
that most code at Databricks is now written by agents, so engineers reach for
lightweight editors (Cursor, VS Code, Neovim) and need fast codebase
orientation — jump-to-definition, workspace symbol search, find-references —
more than exhaustive LSP completion/refactoring coverage. That reframing let
them define a new north-star metric, time-to-initial-intelligence (TTII),
and rebuild three layers of Metals around it: a build-free
mbt content-addressed index, compiler-backed
interactive pipelines for Scala (single
presentation compiler holding 24M lines) and Java (built directly on javac
+ Google's Turbine header compiler), and a
metadata-first BSP integration that queries
285k Bazel targets without putting Bazel on the editor hot path. Metals v2 is
open-sourced under Apache 2.0 and shipping in Cursor, VS Code, and Neovim, with
early external adoption at Stripe.
Key takeaways¶
-
A new metric, TTII, drove the redesign. Time-to-initial-intelligence measures how quickly the editor becomes useful after opening the repo — the clock starts when the language server activates and stops when the most important features (fuzzy workspace symbol search, jump-to-definition, find-references across the repo) are available with no user intervention. Treating startup usefulness as a measurable, improvable metric is the core discipline. (Source: sources/2026-08-11-databricks-open-sourcing-metals-v2)
-
The build server was removed from the startup critical path. Metals v1 waited for a Build Server Protocol (BSP) server to supply the initial project model; Metals v2 indexes workspace sources directly with its own mbt index and owns the initial project model. This is the build-free repo index pattern. (Source: sources/2026-08-11-databricks-open-sourcing-metals-v2)
-
The mbt index is content-addressed against Git.
mbt= Metals Build Tool. It discovers files viagit ls-files --stageand uses Git blob OIDs to decide which index entries can be reused, so incremental updates recompute only changed files (content-hash-incremental-index). Each entry is a file-local summary: package declarations, definitions with source locations, and compact bloom filters of referenced identifiers used to quickly rule out files that cannot contain a reference before doing precise checks. -
The mbt index numbers. In Databricks' monorepo the persisted index is 936MB uncompressed and covers 2.9M symbols across 142k+ Scala/Java/ Protobuf files. A clean benchmark build takes 22s at full utilization across 32 cores; parsing a pre-built index from disk takes 5s. Production TTII (start server, load stale index, update against latest
git ls-files --stage, restart presentation compilers): p50 8.7s, p90 36.7s. Fuzzy symbol search over 2.9M symbols: p50 10ms, p90 95ms. -
A single Scala presentation compiler holds 24M lines in scope. The presentation compiler caches and reuses symbol-table info across runs; combined with lazy symbol resolution, one instance keeps the full 24M-line Scala codebase in scope while publishing diagnostics at p50 0.9s, p90 8.9s and jump-to-definition at p50 7ms, p90 575ms. This is well outside what the Scala compiler was originally built for.
-
Two compiler modes: fallback and precise. Before a build sync a fallback compiler takes a permissive view — every source file in the repo is an eligible dependency candidate — making navigation useful immediately, even across code that doesn't yet compile in Bazel. After a build sync a precise compiler restricts itself to the classpath/sourcepath boundaries the build server reports, giving build-accurate diagnostics (fallback-then-precise-compiler-mode).
-
Outline mode makes a huge sourcepath tractable. Sources not open in the editor are stripped of method bodies before type-checking, preserving type signatures while skipping the method-body work where most type-checking time goes (outline-mode-type-checking). An in-memory source-layout index (built from the mbt index in <50ms) maps classes to source locations so the compiler loads symbols through the index instead of scanning the filesystem.
-
Java is built on
javac+ Turbine, not on JDT/NetBeans. Existing Java language servers are build-centric — the exact coupling v2 removes — so Databricks implemented the Java LSP surface directly onjavacAPIs. Pathologicaljavac"enter"-phase blowups (files transitively importing millions of lines at the symbol-outline layer) are fixed with ajavaSymbolLoader: "turbine-classpath"mode using a modified Turbine header compiler. Turbine processes ~1M lines of Java/sec single-threaded, letting Metals recompile the whole Java codebase on a regular interval and keep cross-file symbols on the classpath; thejavac"analyze" phase then runs near its interactive limit of ~100k lines/sec. -
Metadata-first build integration keeps Bazel off the hot path. Metals v2 still uses BSP but with a narrower contract: routine diagnostics come from the compiler pipelines, while BSP is used mostly to query build metadata — dependencies, generated sources, test discovery, debug launchers (metadata-first-build-integration). The production BSP server is an internal Go implementation tailored to Databricks' Bazel rules, querying 285k Bazel JVM targets at editor latencies.
-
Build sync is an explicit, on-demand user action. In a large monorepo, background IDE sync can take the Bazel lock and compete with developer builds, so sync is never startup behavior or a background task — users sync individual files/directories on demand. Metadata is stored in a JSON snapshot that scales via constant pooling for repeated labels, paths, and repository prefixes; there is deliberately no shared sync configuration format (predefined sync sets grow, get copied between teams, and become slower than focused sync).
Adoption numbers (business case)¶
- By July 2026, 92% of weekly active IDE users open Cursor vs 12% for IntelliJ; among single-IDE engineers, 2.4k on Cursor vs 120 on IntelliJ. Databricks did not renew the majority of IntelliJ seats this year.
- Cursor's share of Scala/Java file-open events rose from 40% to 78% since the first major Metals v2 improvements landed (Oct 2025) — slower than aggregate IDE adoption because IntelliJ was most entrenched for JVM work.
- Prepared with Cursor and core Metals maintainers at VirtusLab; early external adoption at Stripe ("rolling out … less than a month ago … works well in our Java codebase").
Caveats¶
- Active-editing features (refactorings, completions) were deliberately kept narrow in scope — the redesign optimizes navigation/orientation, not full LSP coverage. TTII, not depth of edit features, was the target.
- The production Bazel BSP server is not part of the OSS release; only its design choices are shared.
- The 24M-lines-on-one-presentation-compiler and ~1M-lines/sec Turbine figures are specific to Databricks' hardware and monorepo shape; they illustrate the ceiling the architecture reaches, not a guarantee for other codebases.
Source¶
- Original: https://www.databricks.com/blog/open-sourcing-metals-v2-databricks-java-and-scala-language-server-multi-million-line-codebases
- Raw markdown:
raw/databricks/2026-08-11-open-sourcing-metals-v2-databricks-java-and-scala-language-s-b54d76a4.md
Related¶
- time-to-initial-intelligence
- language-server-protocol
- content-addressed-index
- presentation-compiler
- systems/metals-v2
- systems/google-turbine
- systems/build-server-protocol
- systems/bazel
- concepts/monorepo
- incremental-indexing
- concepts/bloom-filter