Unifying Workers AI and AI Gateway into a single AI control plane¶
Summary¶
Cloudflare describes the convergence of two AI products —
AI Gateway (a proxy to any model provider
with observability, logging, access, and security) and
Workers AI (first-party inference-as-a-service on
Cloudflare-managed GPUs) — into a single AI control plane. The two
products already share entrypoints (the Workers AI binding and a unified
/ai/ REST API); this post lays out the roadmap for treating "connect to a
model" as one path regardless of whether the model runs on Workers AI or an
external provider. The near-term, shipped pieces are automatic observability
for all Workers AI traffic via a default gateway and unified billing (AI
Gateway prepaid credits now spendable on Workers AI). The forward-looking
pieces are model-first routing (specify a model, let the control plane pick
the provider and fail over transparently) and smart routing (a classifier
on Workers AI reads the prompt and a heuristic scorer picks the best model with
zero config).
Key takeaways¶
-
One entrypoint, two products. There is no separate AI Gateway vs Workers AI binding — both go through the same
env.AI.run(model, input, options)path, and a unified/ai/run/...REST endpoint mirrors it. Adding{ gateway: { id: 'default' } }as the third argument routes a Workers AI call through the gateway. (Source: this article) -
defaultgateway = zero-setup observability. Passingdefaultas the gateway ID auto-creates the gateway on the first authenticated request. Every request is then logged with full request/response payloads, per-model token counts, and cost attribution — no dashboard setup. Outgrow it later by creating a named gateway and changing one parameter (e.g. for custom caching rules or per-application traffic splits). This is centralized AI governance made opt-out-by-default. (Source: this article) -
Unified billing now covers Workers AI. Previously AI Gateway credits could only pay external providers (OpenAI, Anthropic, …); now a single prepaid wallet is spendable across OpenAI, Anthropic, Workers AI, or any supported provider — one unified billing surface. Using the unified-billing path on Workers AI also grants elevated rate limits (exact numbers deferred to developer docs). (Source: this article)
-
Model-first routing (coming soon). Provider-first routing forces the caller to reason about infrastructure ("which provider do I call? what if it's down?"). Model-first routing inverts this: you request a model (e.g.
kimi-k2.7-code) and the control plane picks the provider, handles failover, and load-balances. If Workers AI has capacity you get the managed infrastructure; if it's saturated the gateway transparently load-balances to another provider hosting the same weights — no application-level retry or fallback logic. Cloudflare works with "vetted providers" and can respect constraints like Zero Data Retention (ZDR). Piloting "in the coming months." (Source: this article) — see concepts/model-first-routing and model-first-routing-with-transparent-failover. -
Smart routing (piloting internally). Beyond failover: you specify no model and let the gateway decide. A classifier running on Workers AI reads the prompt and predicts task type (coding, research, summarization, general Q&A), complexity, and how much context matters; a heuristic scorer then maps that to the best model from a curated pool. Teams that want control can still pin exact models; the zero-config path aims for better economics and performance without hand-maintained routing logic. (Source: this article) — see patterns/specialized-agent-decomposition.
-
Resiliency becomes a routing problem, not an application problem. With all inference flowing through one central proxy choke point, provider outages/rate-limits are absorbed by the gateway shifting traffic, rather than by retry/fallback code in each Worker. Model availability is treated as a routing input. (Source: this article)
Architecture / mechanics¶
- Binding call routed through the gateway:
const response = await env.AI.run(
'@cf/zai-org/glm-5.2',
{ messages: [{ role: 'user', content: 'Hello!' }] },
{ gateway: { id: 'default' } } // auto-creates the gateway on first use
);
- Unified REST endpoint (non-Workers callers):
curl "https://api.cloudflare.com/client/v4/accounts/{account_id}/ai/run/@cf/zai-org/glm-5.2" \
-H "Authorization: Bearer {api_token}" \
-H "cf-aig-gateway-id: default" \
-H "Content-Type: application/json" \
-d '{ "messages": [{"role":"user","content":"What is the capital of France?"}] }'
- Model-first request (OpenAI-compatible chat completions, model named without a provider):
curl -X POST "https://api.cloudflare.com/client/v4/accounts/{account_id}/ai/v1/chat/completions" \
-H "Authorization: Bearer {api_token}" \
-H "cf-aig-gateway-id: my-gateway" \
-H "Content-Type: application/json" \
-d '{ "model": "kimi-k2.7-code", "messages": [{"role":"user","content":"Review this function"}] }'
model value carries no provider prefix — the gateway resolves it to a
provider (Workers AI, Moonshot's own API, or another host of the same weights).
- Routing evolution ladder disclosed by the post:
- Provider-first (today's default) — model string names the provider.
- Model-first — name the model; gateway selects provider + fails over.
- Smart routing — name nothing; Workers-AI classifier + heuristic scorer pick the model from a curated pool.
Operational numbers / status¶
- Model-first routing: pilot "in the coming months" for all AI Gateway + Workers AI users.
- Smart routing: currently piloting internally; testing/iterating "in the next few weeks before release."
- Unified billing on Workers AI: shipped ("launching today"); elevated rate limits available on the unified-billing path (numbers in developer docs).
- No latency/throughput/accuracy benchmarks for the classifier or scorer are given.
Caveats¶
- Roadmap post, not a system-internals deep-dive. Model-first routing and smart routing are announced-as-coming; the classifier architecture ("a classifier running on Workers AI" + "a heuristic scorer") is described at a conceptual level with no model, feature set, or evaluation numbers.
- Product-launch framing. Included per AGENTS.md borderline rule because the routing-architecture content (control-plane convergence, model-first vs provider-first routing, classifier-based smart routing, resiliency-as-routing) is well over 20% of the body and is genuinely architectural.
- "Vetted providers" and ZDR handling are asserted, not detailed.
Source¶
- Original: https://blog.cloudflare.com/workers-ai-gateway-unification/
- Raw markdown:
raw/cloudflare/2026-08-07-unifying-workers-ai-and-ai-gateway-into-a-single-ai-control-055cea90.md
Related¶
- systems/cloudflare-ai-gateway — the proxy/control-plane tier being unified
- systems/workers-ai — the first-party inference product folded in
- concepts/model-first-routing — name the model, not the provider
- model-first-routing-with-transparent-failover — gateway resolves provider + fails over
- patterns/specialized-agent-decomposition — classifier + heuristic scorer pick the model
- patterns/central-proxy-choke-point — the single choke point that makes routing/governance possible
- patterns/unified-billing-across-providers — one prepaid wallet across providers
- concepts/centralized-ai-governance — observability/cost/policy at one control plane
- companies/cloudflare