Skip to content

DATABRICKS 2026-08-13

Read original ↗

Smart Routing in Unity AI Gateway: Match frontier quality with 30%+ lower cost per task

Summary

Databricks details the architecture of Smart Routing — the AI-Gateway model-router that its earlier cost-management playbook (sources/2026-08-07-databricks-managing-ai-coding-costs-at-scale) named but did not describe — now shipping in Beta inside Unity AI Gateway. The problem: with 33 new models released in 2026 alone and work clustering into capability tiers, defaulting every coding task to the most capable (and most expensive) model wastes money, but asking developers to pick a model per task is choice-overload they resolve by pinning the biggest model and moving on. Smart Routing automatically matches each task to the right model based on complexity. The load-bearing design decision is task-aware routing over per-request routing: because at scale coding-agent cost is dominated by cache hit rate, the router commits a model (and effort level) for the whole session rather than re-deciding per message, which would shatter the prompt cache. Internally, a cheap low-latency classifier labels a task with a handful of semantic fields, and the router triangulates from a default medium model — escalating for frontier-demanding work, delegating down for simple work. Results: Smart Routing beat any single model at 65% of the cost per task of a leading model (Opus 5) on internal workloads, and matched Opus 5 at less than half the cost on public benchmarks. Smart Routing also runs inside Omnigent to route across both models and coding harnesses, and every sub-agent launch goes through the Smart Routing API so sub-agents can get a different harness+model with a fresh cache.

Key takeaways

  1. Task-aware routing is chosen specifically to protect cache hit rate. The post frames two routing approaches and rejects one. Per-request routing (route each message by its prompt's complexity) loses because "at scale, costs are dominated by cache hit rate" and a high hit rate requires routing consecutive turns to the same model and effort level. Task-aware routing assesses complexity once at session start, sticks with that model+harness for the session, preserves the cache, and leaves room to upgrade/downgrade later when the cache goes stale anyway (e.g. at a compaction event). (concepts/model-first-routing, Source: sources/2026-08-13-databricks-smart-routing-in-unity-ai-gateway)

  2. Most of the win is using cheaper models for simpler tasks. Internal task mix is diverse and most tasks don't need a premium model; the router is configured to pick cheaper models for simple tasks and "escalate" complex work that needs frontier-model performance — the operationalization of cheapest-capable model routing.

  3. A cheap classifier gates the expensive model (two-stage router). A cheaper, low-latency model reads the task description and labels it with semantic fields — what part of the system changes, what code evidence the prompt carries (a snippet, a traceback, or nothing), how it appears to be failing, how localized the fix looks, what kind of project it belongs to — from which the router derives a task-type family and language family. Using a frontier model as the extractor would tax every request (even the cheap ones you're trying to save on), so the extractor is intentionally small and fast. (patterns/two-stage-evaluation)

  4. The router triangulates from a medium default. It defaults to a medium-sized model and moves in either direction from the labels — escalating to a more expensive model when the task demands frontier capability/knowledge, delegating down to a cheaper one when it doesn't. A single policy applied uniformly to every task lets one router leverage a whole suite of models.

  5. Routing used only start-of-task information and still worked. They gave the router only the task description and metadata — no answer, tests, or repository context — and it still chose effectively. This is what makes task-aware routing viable: the decision is made before the task begins, on the client side.

  6. Quantified results. Internal benchmark (no lab had access): 35% savings; public coding benchmarks (to show generalization): 56% savings. Headline framings: Smart Routing outperformed any single model at 65% of the cost per task of Opus 5 internally, and matched Opus 5 at <50% cost on public benchmarks. Marketing line: 30%+ lower cost per task at matched frontier quality. A router with perfect foresight would beat every single model at a fraction of current spend — so there's substantial headroom.

  7. Smart Routing spans models AND harnesses via Omnigent. In Omnigent, developers can pick "Smart Routing" instead of a specific harness; Omnigent then selects both harness and model per task, with model routing powered by Smart Routing in Unity AI Gateway. This is the meta-harness task-routing pattern with the model-selection half now concretely specified.

  8. All sub-agent launches route through the Smart Routing API. The user's initial prompt is often underspecified and hard to size, so sub-agents let the system re-route new work on new information — a fresh cache and clear instructions. A single task can experience nuanced routing across planning and parallel sub-agents (e.g. route large-codebase summarization to cheaper models while designing the architecture with more expensive ones), compounding savings.

  9. Feedback signal: log every session trace, evaluate for cost AND productivity. Routers are early tech needing iteration, so the first step was logging all coding-session traces (recorded into Unity Catalog, governed by tagging + access policies because trace data is highly sensitive). AI + human review of traces evaluate router changes; monitored metrics: session breakdown by model, # sessions completed end-to-end by routed models, and $ savings. They explicitly do not optimize cost at the expense of developer productivity.

  10. Where it's going: route where scoping is free, or route after a few turns. Fully-specified, machine-authored task statements (PR reviews, sub-agent launches, batch migrations, scheduled jobs) route well today — which is why Databricks is deploying there first. For interactive work the opening prompt is the worst moment to decide; a cheap model can handle the initial exchange and ask clarifying questions, then route once the task takes shape. Making mid-session switching affordable requires the routing layer to explicitly price cache misses and use context compaction as the natural switch seam (a cache miss is already happening there — "this is how Cognition's Devin Fusion does it"). (route-at-context-compaction-seam)

Systems / concepts / patterns extracted

Operational numbers

  • 33 new models released in 2026 (as of the post).
  • Internal benchmark savings: 35%. Public-benchmark savings: 56%.
  • Smart Routing = 65% of the cost per task of Opus 5 (internal), while outperforming any single model.
  • Matched Opus 5 at <50% cost on public benchmarks; 30%+ lower cost/task headline.
  • Prior post's companion figures: harness/cache tuning ≈ 50% token reduction; Smart Router ">30% avg task-cost reduction at matched quality."

Caveats

  • Tier-3 source + Beta product launch. This is a Databricks product announcement, but it clears the borderline-include bar: it discloses genuine serving-infra architecture (task-aware vs per-request routing, cache-hit-rate economics, a two-stage classifier→router pipeline, sub-agent routing through a gateway API, compaction-seam switching) well above the 20% threshold.
  • Classifier is described conceptually — no model name, feature weights, or classifier accuracy/latency numbers.
  • The router uses "a single policy" today (one policy for all tasks); per-task policy differentiation is future work.
  • Benchmark tasks are "unusually well-behaved" self-contained statements of work; Databricks explicitly flags that routers do worse on messy real interactive sessions, and that opening prompts are a poor routing signal — the honest limitation the roadmap addresses.
  • "Perfect-foresight router" comparison is a thought experiment, not a shipped result; it quantifies headroom, not current performance.

Source

Last updated · 766 distilled / 2,225 read