Cut your AI spend with AI Gateway's Auto Router¶
Summary¶
Cloudflare released Auto Router in public beta through AI Gateway:
set the model to cloudflare/auto and the gateway routes each request to a model
that is "capable enough for the task" without the user picking a model. It is the
production realization of the smart-routing roadmap item Cloudflare previewed
when it unified Workers AI and AI Gateway into one control
plane. Internally (via Cloudflare's OpenCode harness and Cloudflare OS),
it reports up to 30% cost savings vs. frontier-only usage. The mechanism: a
multi-head classifier on Workers AI scores each request across 14 task categories
plus four difficulty dimensions; a benchmark-derived scoring matrix estimates each
candidate model's fit; then the router picks the model maximizing
utility = expected quality − adaptive cost penalty, where the cost penalty also
accounts for KV-cache read/write cost of switching models mid-session.
Key takeaways¶
-
cloudflare/autois a model alias that routes, not a model. The caller sets one model string and the gateway resolves the actual model per request — a direct extension of model-first routing (name a capability, let the control plane pick) into zero-config model selection. (Source: sources/2026-09-30-cloudflare-cut-your-ai-spend-with-ai-gateways-auto-router) -
Two-stage architecture: classifier → scoring matrix. A multi-head classification model runs on Workers AI on edge GPUs. It emits (1) probabilities across 14 task categories (coding, planning, research, data analysis, …) and (2) ratings on a 1–5 scale across four dimensions: complexity, ambiguity, stakes, and dependence on earlier context. A separate scoring matrix combines those signals with per-model benchmark results to estimate fit. "Routing decisions are legible because you can inspect each task's predicted category and complexity to see how it translated into the model choice."
-
Utility function trades quality against price, adaptively.
utility = expected quality − adaptive cost penalty. On easy requests the price term dominates, so a smaller capable model can win; as difficulty rises the cost penalty shrinks and stronger models have room to win. This is the don't-pay-frontier-rates-for-non-frontier-work economics applied by a classifier rather than a calibrated-uncertainty gate. -
Cache-aware switching penalty is the novel cost model. For long agentic sessions, cost is driven less by list price than by KV-cache read cost, which grows with context length. Switching models throws away the hot cache and forces the new model to re-write the whole context. So: within a turn the cache is hot and switching rarely pays off (keep the same model); across turns a switching penalty grows with the number of tokens already in context — a model still holding a live cache is priced at its cheap cache-read rate; every other candidate is priced at the full context-rewrite cost. "The deeper the conversation, the more a switch has to earn back." Reasoning tokens compound this: most models can't read another model's reasoning tokens, so a switch that drops them may force the new model to redo that work at output prices. Roadmap: prefer staying within the same model family on a switch.
-
"Jagged frontier" framing. "The ability to solve a problem often exists somewhere in this portfolio of models; the router's job is to choose the right model for each task while balancing quality and price." Savings come from not paying frontier rates for non-frontier work and grow with how much non-frontier work you have (see concepts/efficiency-frontier).
-
Lower per-token price ≠ lower total cost. "A model that looks cheaper on paper may end up using disproportionately more tokens to solve a problem." The router minimizes predicted trajectory cost, not dollars-per-million-tokens load balance.
-
Pool construction is a filter chain, then a ranked fallback. For each request the gateway first builds the pool of models that can serve it — filtering on request format / execution mode, plus attached credentials, billing config, access-control policies, and spend limits, and dropping unhealthy upstreams (auto-returned after outage). It then returns a ranked list; AI Gateway tries the winner and can fall through to another eligible model if a provider can't serve it — the same provider-abstraction + automatic-failover primitive AI Gateway already had, now driven by the router's ranking.
-
No retraining on new models. Adjusting the router for a newly released model only requires adding its benchmark-derived weights to the scoring matrix — not retraining the classifier. The same classifier can back multiple routing profiles; a planned
cloudflare/auto-bestreuses the classification + pool but selects highest expected quality with no cost tradeoff.
Operational numbers¶
Internal "general knowledge work" benchmark — 97 tasks × 3 samples/model (291 trials), simulated workspace tools (email, calendar, Slack, files, travel, finance); each task requires tool use to produce a verifiable answer/action:
| Model | Success rate | Total cost | Cost per success |
|---|---|---|---|
cloudflare/auto |
86.6% (+6.2/−6.9 pp) | $2.10 | $0.0084 |
| Anthropic Claude Opus 5.5 | 96.6% (+2.7/−3.8 pp) | $5.91 | $0.0210 |
| OpenAI GPT-6 Sol | 84.2% (+6.5/−6.9 pp) | $2.64 | $0.0108 |
- Auto Router came in at ~80% the cost of Sol and ~35% the cost of Opus at comparable-to-slightly-lower quality.
- Reported up to 30% cost savings in internal coding use via OpenCode vs. frontier-only.
- Confidence intervals are 95% CIs from 10,000 task-level bootstrap resamples (see concepts/benchmark-methodology-bias).
- Free during beta.
Roadmap / what's next¶
Expand the cloudflare/auto model pool; add zero-data-retention requirements
to the model filter; account for provider capacity in selection; select the
appropriate reasoning/thinking level per request; full Responses API + WebSockets
support; explore structured decision models as a first-pass classifier;
cloudflare/auto-best profile; prefer same-family switches to preserve reasoning
tokens.
Caveats / what the source does not disclose¶
- Quality is lower than Opus (86.6% vs 96.6%) — Auto Router trades ~10 pp of success rate for ~64% cost reduction on this benchmark; whether that trade is right is workload-dependent.
- The benchmark is Cloudflare's internal general-knowledge-work benchmark, not a public/third-party one — treat cross-model comparisons with the usual first-party-benchmark caution.
- The classifier's architecture, training data, task-category taxonomy, and the exact scoring-matrix weights are not published; the "legibility" claim is asserted, not demonstrated with an example trace.
- The default
cloudflare/automodel pool ("default models") is referenced via docs, not enumerated in the post. - "Up to 30%" is a ceiling from internal coding usage, not a guaranteed or average saving.
Source¶
- Original: https://blog.cloudflare.com/auto-router/
- Raw markdown:
raw/cloudflare/2026-09-30-cut-your-ai-spend-with-ai-gateways-auto-router-ac0f8886.md
Related¶
- systems/cloudflare-ai-gateway — Auto Router ships as a feature of AI Gateway.
- systems/workers-ai — hosts the multi-head classifier on edge GPUs.
- systems/cloudflare-os — internal consumer of
cloudflare/auto. - concepts/model-first-routing — the routing model Auto Router realizes as zero-config selection.
- concepts/efficiency-frontier — the "jagged frontier" quality/price curve.
- concepts/kv-cache — the switching-penalty cost driver.
- patterns/cheap-approximator-with-expensive-fallback — cheap-for-easy / expensive-for-hard economics.
- patterns/ai-gateway-provider-abstraction — the pool/failover primitive Auto Router rides on.
- companies/cloudflare