Skip to content

SYSTEM Cited by 1 source

Ultron (history-based query optimization)

Ultron is Databricks' history-based query optimization framework. It improves optimizer decisions — such as which join operator to use — by leveraging the repetitive nature of analytical workloads: the same or similar queries run over and over, so the outcomes of past executions are a strong signal for optimizing future ones. Published as a VLDB 2026 paper (presenter: Eric Liang). (Source: sources/2026-08-27-databricks-building-for-the-ai-era-lakebase-streaming-and-lakehouse-innovations-vldb-2026)

Motivation

Lakehouse query latency could be significantly improved "if only the optimizer had near-perfect knowledge about the data." Traditional cost-based optimizers rely on statistics that are often stale or missing. Ultron's insight is that execution history is a cheaper, more accurate source of that knowledge for recurring analytical workloads than up-front statistics collection.

How it works

  1. Records execution history. Ultron efficiently stores the history of executed queries and manages the logs of those executions.
  2. Feeds the optimizer. On subsequent (repeated or similar) queries, the optimizer uses that history to make better choices — e.g. selecting the type of join operator — approximating the "near-perfect knowledge" ideal.

The two named engineering challenges are (a) storing query history efficiently and (b) managing the logs of executed queries at production scale.

Results

  • Improved median join latency by 25% on production workloads. (Vendor-reported.)

Relationship to other systems

  • Complements Adaptive Query Execution (which adapts within a single query's runtime using observed statistics); Ultron instead learns across query executions over time.
  • Related to the join-order agent line of work on making better join decisions, and to optimizer statistics as the classical alternative signal.

Seen in

Last updated · 766 distilled / 2,225 read