Skip to content

CONCEPT Cited by 6 sources

Model-first routing

Definition

Model-first routing is an inference-routing model where the caller specifies which model they want — a capability, e.g. "a capable reasoning model" or a named model like kimi-k2.7-code — and a control plane decides which provider actually serves it. It is the inverse of provider-first routing, where the caller must know and name the provider that hosts a model and is responsible for retry/fallback when that provider is down or rate-limiting.

The distinction is about what the caller has to reason about:

  • Provider-first: "Which provider do I call? What if they're down?" Infrastructure concerns leak into application code.
  • Model-first: "What do I need — a reasoning model? a cheap embedding model?" The control plane owns provider selection, failover, and load balancing.

(Source: sources/2026-08-07-cloudflare-unifying-workers-ai-and-ai-gateway)

Why it matters

When many hosts serve the same model weights (a first-party managed platform plus the model author's own API plus other vetted hosts), the provider choice stops being a semantic decision and becomes a routing decision — best made centrally where capacity, health, and policy are visible. This turns resiliency into a routing problem rather than an application problem: a provider outage or rate-limit is absorbed by the router shifting traffic to another host of the same weights, with no application-level retry logic or fallback chains in the caller (model-first-routing-with-transparent-failover).

It also unlocks zero-config model selection as a further step: if the router can pick a provider for a named model, a classifier can pick the model itself from a curated pool based on the prompt (patterns/specialized-agent-decomposition, i.e. "smart routing").

Preconditions

  • A central proxy all inference flows through, so the router sees every request and every provider's health.
  • Weight-equivalence across providers — the router can only substitute providers that host the same model, so output quality is preserved. (This is why Cloudflare emphasizes "vetted providers" and per-request constraints like Zero Data Retention.)
  • A model namespace decoupled from the provider — in Cloudflare's API the model field carries no provider prefix ("model": "kimi-k2.7-code"), which is what makes it provider-agnostic. Compare provider-first calls where the provider selector lives inside the model string (@cf/…, @provider/model). This is a form of protocol-normalization at the model-identifier layer.

Seen in

  • systems/cloudflare-ai-gateway — cloudflare/auto Auto Router (public beta, 2026-09-30): the shipped realization of zero-config model selection. The smart-routing roadmap item below became a product. A caller sets one model alias (cloudflare/auto) and the gateway picks the model per request via a two-stage, legible pipeline: a multi-head classifier on Workers AI (14 task categories + 1–5 ratings on complexity/ambiguity/stakes/context-dependence), a benchmark-derived scoring matrix (new models added by weights, no retraining), and a utility function utility = expected quality − adaptive cost penalty (price dominates on easy requests; the penalty shrinks as difficulty rises). This is the routing key generalized from provider (failover) and compliance (Concurrence) and static task tier (Databricks) to a classifier-scored quality-vs-cost tradeoff, applied at the central choke point. Two properties worth carrying forward: (1) the router minimizes predicted trajectory cost, not per-token list price — a cheaper-per-token model that burns more tokens can lose; (2) it adds a cache-aware switching penalty — across turns, switching models is priced at the full context-rewrite cost while the incumbent keeps its cheap cache-read rate, so "the deeper the conversation, the more a switch has to earn back." Reported ~35% of Opus cost at 86.6% vs 96.6% success on Cloudflare's internal benchmark; up to 30% coding savings. A planned cloudflare/auto-best reuses the same classifier + pool but drops the cost term. (Source: sources/2026-09-30-cloudflare-cut-your-ai-spend-with-ai-gateways-auto-router)

  • User Insights — where the routing signals become observable to a human first. The 2026-09-30 User Insights update surfaces the same task-analysis signal (task category + complexity/ambiguity/stakes/ context-dependence) that the Auto Router consumes, but as a read-only model overkill / task / turns view — letting a team see which users/agents send simple tasks to over-capable models before any automated routing kicks in. The signal is produced by an async classification Worker off the request path (not inline), so model-first routing here is split into an offline observe half (User Insights) and an inline act half (Auto Router) sharing one classifier. (Source: sources/2026-09-30-cloudflare-identify-ai-model-overuse-with-user-insights)

  • systems/cloudflare-ai-gateway — model-first routing announced as the next step after unifying Workers AI and AI Gateway into one control plane; if Workers AI is at capacity the gateway transparently load-balances a named model to another provider hosting the same weights. (Source: sources/2026-08-07-cloudflare-unifying-workers-ai-and-ai-gateway)

  • systems/concurrence / systems/unity-ai-gateway — compliance-gated variant. At Concurrence, model routing is "compliance-gated before it is cost-gated": eligibility is filtered on compliance (HIPAA / BAA coverage) before capability, quality, or cost. The enforcement lever is a per-workload endpoint resolver restricted to the BAA-covered namespace, so PHI can't be routed to an uncovered model; Smart Routing then optimises only within that boundary. Adds a compliance precondition to the "weight-equivalence" precondition below — the router substitutes only among compliant hosts. (Source: sources/2026-09-23-databricks-how-concurrence-governs-clinical-ai-at-a-trillion-token-scal)

  • systems/databricks-foundation-model-api — task-difficulty tiering variant. A Databricks security-review system routes by task complexity across a fixed model ladder rather than by provider capacity or weight-equivalence: "Haiku handles lightweight classification, Sonnet handles most review work, and Opus is reserved for the heaviest reasoning." This is the "pick the right model for the job" idea applied along the cheap→capable axis (patterns/cheap-approximator-with-expensive-fallback), a distinct routing key from the provider-abstraction / failover framing above. (Source: sources/2026-09-24-databricks-how-i-built-agent-based-security-reviews-on-databricks)

  • systems/unity-ai-gateway — promotion-to-router-toolkit variant. In Databricks' Day-1 model rollout (sources/2026-09-28-databricks-how-databricks-rolls-out-frontier-models-to-12000-employees), a newly-evaluated model that is cheaper but not higher-quality than the incumbent isn't made the default — it's added to the smart router's toolkit so the router can pick it for tasks where its cost advantage wins. GPT-6 Sol, which was occasionally a quality downgrade from GPT-5.6 Sol but ~50% cheaper, took exactly this path (not the Codex default, but in the router pool). Shows the router pool as a promotion outcome of the model release lifecycle, not just a static configuration. (Source: sources/2026-09-28-databricks-how-databricks-rolls-out-frontier-models-to-12000-employees)

Merged aliases

  • llm-workflow-router- dynamic-model-selection-for-agents
  • task-aware-vs-per-request-routing
  • compliance-gated-routing
Last updated · 766 distilled / 2,225 read