Skip to content

SYSTEM Cited by 8 sources

Unity AI Gateway (Databricks)

Unity AI Gateway is Databricks' productised instance of the AI-gateway provider abstraction pattern, specialised to coding agents + MCP integrations rather than just application LLM calls. Its job is to be the single governance + cost + telemetry plane for all coding-tool traffic in a Databricks customer's fleet.

Generalisation to org-wide agent populations (2026-05-20 disclosure)

The 2026-05-20 Governing AI agents at scale with Unity Catalog post (sources/2026-05-20-databricks-governing-ai-agents-at-scale-with-unity-catalog) generalises Unity AI Gateway from coding-agent scope to org-wide agent scope and adds three named architectural extensions that this page didn't previously canonicalise.

Scope generalisation

The 2026-04-17 launch positioned the Gateway around coding-agent sprawl (Cursor / Codex / Claude Code / Gemini CLI). The 2026-05-20 post generalises to every department's agents: dev (coding agents), analytics (forecasting agents), sales-ops (lead-scoring agents), support (ticket-routing agents), marketing (personalization), finance (reconciliation). The architectural surface didn't change — the "every model call, every tool invocation, every agent interaction flows through the gateway" principle now covers all of them.

Three new feature surfaces

The Gateway as disclosed on 2026-04-17 had centralised audit + cost + observability. The 2026-05-20 post discloses three additional named layers attached to the same proxy:

Layer What it does Wiki entity
Service Policies Pre-execution per-tool-call evaluation; UC functions attached to registered MCPs; returns allow/deny/consent; fail-closed on deny systems/uc-service-policies
Guardrails Inline content scanning of every model call — inputs (PII, jailbreak), outputs (hallucinations, sensitive content); fail-closed systems/unity-ai-gateway-guardrails
Inference Tables Full payload of every model call (exact prompt + exact response + tokens + latency) written to UC-managed Delta tables; customer-controlled retention systems/inference-tables
Budgets Per-user / per-group monthly spend thresholds with alerts; hard enforcement on roadmap systems/unity-ai-gateway-budgets

Four-pillar repositioning

The 2026-05-20 post repositions the Gateway in the four-pillar framing (concepts/centralized-ai-governance):

  • Pillar 1 (Delegated access) — three-layer composition: OBO permissions + Service Policies + Guardrails. The Gateway is the enforcement fabric where all three layers run.
  • Pillar 2 (Data-centric AI governance) — Gateway writes Inference Tables + UC audit logs to the lakehouse, joinable with business data; substrate for Lakewatch (agentic SIEM).
  • Pillar 3 (Cost intelligence) — usage-tracking + Budgets.
  • Pillar 4 (Open and interoperable) — single governed endpoint across Databricks-hosted models + Azure OpenAI + AWS Bedrock + Anthropic; framework-agnostic across LangGraph / CrewAI / OpenAI SDK / Anthropic SDK / AutoGen / LlamaIndex.

Identity propagation (now explicit)

The 2026-05-20 post is the first to explicitly disclose OBO as the data-access mechanism: "identity flows end to end, from the user who asks the question to the specific table row the agent retrieves." The Gateway is the identity-translation point — agents inherit the invoking user's UC permissions in real time via on-behalf-of token passing, not via shared service accounts.

Three-pillar architecture (from the 2026-04-17 launch post)

  1. Centralised security and audit.
  2. Every agent data-access flow logged in Unity Catalog (same governance substrate as Lakehouse data + ML assets).
  3. All tracing in MLflow (specifically MLflow 3 GenAI tracing — named for Claude Code integration).
  4. MCP servers "managed in Databricks" — the gateway is the policy point for MCP traffic, not just LLM traffic.
  5. Single-identity plane: developers authenticate once with Databricks credentials for all tools (GitHub, Atlassian, etc.), "no separate logins per service".
  6. Single bill and cost limits.
  7. Foundation Model API provides first-party inference for OpenAI, Anthropic, Gemini, and open models like Qwen.
  8. Admins can also "bring external capacity in", extending governance "to all your tokens, regardless of where they flow" — patterns/unified-billing-across-providers.
  9. Gateway-enforced budgets are per-developer, not per-tool — admins give each developer one budget and the developer burns it on whichever tool of choice (Cursor / Codex / Gemini CLI / Claude Code / …).
  10. Full observability in the Lakehouse.
  11. Coding-tool metrics + traces land in Unity-Catalog-managed Delta tables via OpenTelemetry ingestion.
  12. Joinable with other Lakehouse datasets (Workday for adoption-by-org / region / seniority; PR-cycle data for velocity quantification) — patterns/telemetry-to-lakehouse.
  13. Surfaces rate-limit hits as a proactive capacity-planning signal.

Supported clients (at launch)

Relation to existing wiki entities

What the post does not disclose

  • Gateway internals: routing, fallback, rate-limiter algorithm, streaming handling, per-provider adapter shape.
  • MCP-governance mechanics: how the gateway inspects MCP traffic, auth flow from coding-tool → gateway → MCP → data source.
  • Telemetry schema landing in Delta tables.
  • Latency / throughput / cost-per-token / adoption numbers.

Tier-3 Databricks post — ingested because the problem framing (coding-agent sprawl) and three-pillar architecture are substantive, not because the internals are disclosed.

Smart Routing — the model-router (2026-08-13 disclosure)

The 2026-08-13 Smart Routing post (sources/2026-08-13-databricks-smart-routing-in-unity-ai-gateway) discloses the architecture of the Gateway's model router — the "AI Gateway Smart Router" that the 2026-08-07 cost playbook only named. Smart Routing is now in Beta and works directly in Claude Code and Codex.

  • Task-aware, not per-request. The load-bearing decision is task-aware routing: the router picks a model + effort level once at session start and holds it for the session, because at scale coding-agent cost is dominated by cache hit rate and switching models mid-session breaks the prompt cache. Per-request routing is explicitly rejected on cache-cost grounds.
  • Two-stage classifier → router. A cheap low-latency classifier labels the task (system area, code evidence, failure mode, fix localization, project type) → task-type + language families; the router then triangulates from a medium default, escalating for frontier-demanding work and delegating down for simple work, under a single policy (patterns/two-stage-evaluation). It routes on start-of-task info only (description + metadata).
  • Spans models AND harnesses via Omnigent. In Omnigent, "Smart Routing" is selectable in place of a fixed harness; Omnigent picks harness + model per task, model routing powered by Smart Routing (patterns/specialized-agent-decomposition). Every sub-agent launch goes through the Smart Routing API, so sub-agents get a different harness+model with a fresh cache.
  • Trace feedback loop. All coding-session traces are logged into Unity Catalog (governed by tagging + access policy, as trace data is sensitive) and reviewed by AI + humans to iterate the router; monitored metrics are session-by-model breakdown, # end-to-end-completed-by-routed-model, and $ savings.
  • Roadmap: route where scoping is free first (PR reviews, sub-agent launches, batch/scheduled jobs); route interactive work after a few turns; and switch models at the context-compaction seam by explicitly pricing cache misses.
  • Results: 35% (internal) / 56% (public-benchmark) cost savings; outperformed any single model at 65% of Opus 5's cost per task; matched Opus 5 at <50% cost; "30%+ lower cost per task" headline.

Still undisclosed: classifier model/features and its accuracy/latency; the routing policy's exact thresholds; per-task (vs single) policy.

Seen in

The trace table as a tool-failure diagnostic substrate (2026-09-01)

The 2026-09-01 "How we eliminated $1M/year of wasted AI agent spend in one hour" post (sources/2026-09-01-databricks-how-we-eliminated-1-million-a-year-of-wasted-ai-agent-spend) puts the Gateway's OpenTelemetry tracing to a new use: diagnosing silent tool-server failures. Because the Gateway sits on the path of every call, it automatically emits one OTel trace per MCP tool invocation — tool name, arguments, error (if any), token counts, latency, and a session ID tying calls together — into a single unified trace table (now in Beta). No new instrumentation is required; the data is a byproduct of routing.

Pointing Genie One at that table let Databricks ask, in plain English, which tool errors recur most, how many turns each takes to recover, and what each costs in tokens and wall-clock. That surfaced seven silent tool bugs costing an estimated $499K/year in tokens + ~12,000 engineering hours/year (~$1.2M/year) — found, quantified, and fixed in about an hour (trace-driven-tool-failure-diagnosis). This is the same telemetry-to-lakehouse shape the Gateway already used for cost/audit, now aimed at silent tool-failure cost rather than model spend.

The Gateway as the Day-1 model rollout plane (2026-09-28)

The 2026-09-28 How Databricks rolls out frontier models to 12,000 employees on Day 1 post (sources/2026-09-28-databricks-how-databricks-rolls-out-frontier-models-to-12000-employees) is the first ingested source to use the Gateway as the control plane for a model release lifecycle — the pipeline that gives 12,000+ employees Day-1 access to newly-released models while assessing whether they're actually good long-term workhorses (patterns/experimental-tier-model-promotion). It adds two mechanisms the wiki hadn't canonicalised on this page.

Server-side designation + laptop-side config push (the UG CLI half)

Server-side config is not enough: engineers run Claude Code, Codex, and Omnigent locally, so the Gateway enabling a model server-side has to reach every laptop. The UG CLI — deployed via Mobile Device Management and already running on everyone's machine — runs whenever a harness starts, checks the Gateway for new models / tools / skills, and updates the local harness config. So Unity Gateway is both (a) the place where a model is switched on and designated default vs experimental, prepared for smart routing, and trace-collected; and (b) via the UG CLI, the fan-out that makes "Day 1" literal for a laptop fleet (patterns/progressive-configuration-rollout). Experimental models surface in the harness with an "Experimental" tag so users can select them but understand they may not stick around.

Four per-user budgets (adds quality-frontier + experimental)

The post extends Budgets from the dual (daily + monthly) architecture to four budgets, adding two scoped to a model class: a quality-frontier budget (rations the most premium models — GPT Astra, Claude Fable — used for specialized tasks at a 2–3× premium, not as daily drivers) and an experimental budget (bounds Day-1 exposure of new, untested models). A new model is tagged into the experimental bucket, so exposing it fleet-wide has bounded cost during the ~3-day evaluation window. See systems/unity-ai-gateway-budgets.

Promote/drop on three signals, with OTel cost normalisation

The promote/drop decision fuses private benchmarks (offline suites — document reasoning, workspace search, Genie — plus online side-by-side runs on real PR creation), user-reported quality (power users over Slack + surveys), and cost tracking from the Gateway's OTel traces. Because early adopters try harder problems with new models, a raw $/session average is misleading — Databricks stratified sessions (single- vs multi-turn × file-edits-or-not) and re-weighted before comparing (a cost-side sibling of patterns/stratified-evaluation-sampling), yielding Opus 4.8→5.5 −29% and GPT-5.6→6 Sol −48% per session. Outcomes vary: Opus 5.5 → GA + slated default for Claude Code; GPT-6 Sol → GA but not default, only added to the smart router's toolkit for its cost advantage — a concepts/model-first-routing outcome. Public benchmarks are treated as insufficient (concepts/benchmark-methodology-bias).

Compliance-gated routing in production — Concurrence (2026-09-23)

The 2026-09-23 How Concurrence governs clinical AI at a trillion-token scale with Unity Gateway post (sources/2026-09-23-databricks-how-concurrence-governs-clinical-ai-at-a-trillion-token-scal) is the first ingested source where a customer runs the Gateway as the intended common inference control point across production, batch, and developer AI at scale — and it adds a governance ordering the wiki hadn't canonicalised: compliance-gated routing. At Concurrence, a clinical-AI company, "routing is compliance-gated before it is cost-gated" — a model must clear a workload's compliance requirements (HIPAA / BAA coverage) before quality, performance, or cost are considered. This inverts the usual Smart-Routing optimisation: the Gateway's model routing and Smart Routing optimise only within each workload's compliance boundary, not across all available models. See concepts/model-first-routing and concepts/centralized-ai-governance.

Concrete mechanisms disclosed:

  • ug coding CLI at scale. All coding-agent model + tool traffic routes through the Gateway's ug coding CLI; each request stays tied to the invoking engineer's identity (patterns/on-behalf-of-agent-authorization), MCP access is centrally managed with per-group permissions and per-user auth to Databricks / Datadog / Linear. July: 14 users, 35.85B input tokens, 95.37% cache reads (concepts/cache-hit-rate); since ug rolled out July 10, ~360k requests, ~61B cumulative input tokens. Attribution via ug usage (individual) and system.ai_gateway.usage (org-level) — patterns/telemetry-to-lakehouse.
  • BAA-namespace endpoint resolver. For batch ai_query on Databricks-hosted Claude, the endpoint resolver only permits models within the BAA-covered namespace, structurally preventing PHI from reaching an uncovered model — the same covered path also runs safety classification (self-harm, suicidal ideation, medical emergencies).
  • Feature-flagged real-time path with a synthetic canary. Real-time patient/clinician inference is built and feature-flagged behind the Gateway with a synthetic canary continuously testing it end to end; production traffic moves as compliance coverage lands.
  • Smart Routing under a compliance boundary. Concurrence is testing the Gateway's Smart Routing against its own healthcare-specific routing, with model eligibility gated on compliance first.

Caveat: Databricks customer story — the only concrete routing lever disclosed is the BAA-namespace resolver; no router internals, no per-workflow latency.

Last updated · 766 distilled / 2,225 read