Skip to content

CLOUDFLARE 2026-08-07 Tier 1

Read original ↗

Unifying Workers AI and AI Gateway into a single AI control plane

Summary

Cloudflare describes the convergence of two AI products — AI Gateway (a proxy to any model provider with observability, logging, access, and security) and Workers AI (first-party inference-as-a-service on Cloudflare-managed GPUs) — into a single AI control plane. The two products already share entrypoints (the Workers AI binding and a unified /ai/ REST API); this post lays out the roadmap for treating "connect to a model" as one path regardless of whether the model runs on Workers AI or an external provider. The near-term, shipped pieces are automatic observability for all Workers AI traffic via a default gateway and unified billing (AI Gateway prepaid credits now spendable on Workers AI). The forward-looking pieces are model-first routing (specify a model, let the control plane pick the provider and fail over transparently) and smart routing (a classifier on Workers AI reads the prompt and a heuristic scorer picks the best model with zero config).

Key takeaways

  1. One entrypoint, two products. There is no separate AI Gateway vs Workers AI binding — both go through the same env.AI.run(model, input, options) path, and a unified /ai/run/... REST endpoint mirrors it. Adding { gateway: { id: 'default' } } as the third argument routes a Workers AI call through the gateway. (Source: this article)

  2. default gateway = zero-setup observability. Passing default as the gateway ID auto-creates the gateway on the first authenticated request. Every request is then logged with full request/response payloads, per-model token counts, and cost attribution — no dashboard setup. Outgrow it later by creating a named gateway and changing one parameter (e.g. for custom caching rules or per-application traffic splits). This is centralized AI governance made opt-out-by-default. (Source: this article)

  3. Unified billing now covers Workers AI. Previously AI Gateway credits could only pay external providers (OpenAI, Anthropic, …); now a single prepaid wallet is spendable across OpenAI, Anthropic, Workers AI, or any supported provider — one unified billing surface. Using the unified-billing path on Workers AI also grants elevated rate limits (exact numbers deferred to developer docs). (Source: this article)

  4. Model-first routing (coming soon). Provider-first routing forces the caller to reason about infrastructure ("which provider do I call? what if it's down?"). Model-first routing inverts this: you request a model (e.g. kimi-k2.7-code) and the control plane picks the provider, handles failover, and load-balances. If Workers AI has capacity you get the managed infrastructure; if it's saturated the gateway transparently load-balances to another provider hosting the same weights — no application-level retry or fallback logic. Cloudflare works with "vetted providers" and can respect constraints like Zero Data Retention (ZDR). Piloting "in the coming months." (Source: this article) — see concepts/model-first-routing and model-first-routing-with-transparent-failover.

  5. Smart routing (piloting internally). Beyond failover: you specify no model and let the gateway decide. A classifier running on Workers AI reads the prompt and predicts task type (coding, research, summarization, general Q&A), complexity, and how much context matters; a heuristic scorer then maps that to the best model from a curated pool. Teams that want control can still pin exact models; the zero-config path aims for better economics and performance without hand-maintained routing logic. (Source: this article) — see patterns/specialized-agent-decomposition.

  6. Resiliency becomes a routing problem, not an application problem. With all inference flowing through one central proxy choke point, provider outages/rate-limits are absorbed by the gateway shifting traffic, rather than by retry/fallback code in each Worker. Model availability is treated as a routing input. (Source: this article)

Architecture / mechanics

  • Binding call routed through the gateway:
const response = await env.AI.run(
  '@cf/zai-org/glm-5.2',
  { messages: [{ role: 'user', content: 'Hello!' }] },
  { gateway: { id: 'default' } } // auto-creates the gateway on first use
);
  • Unified REST endpoint (non-Workers callers):
curl "https://api.cloudflare.com/client/v4/accounts/{account_id}/ai/run/@cf/zai-org/glm-5.2" \
  -H "Authorization: Bearer {api_token}" \
  -H "cf-aig-gateway-id: default" \
  -H "Content-Type: application/json" \
  -d '{ "messages": [{"role":"user","content":"What is the capital of France?"}] }'
  • Model-first request (OpenAI-compatible chat completions, model named without a provider):

curl -X POST "https://api.cloudflare.com/client/v4/accounts/{account_id}/ai/v1/chat/completions" \
  -H "Authorization: Bearer {api_token}" \
  -H "cf-aig-gateway-id: my-gateway" \
  -H "Content-Type: application/json" \
  -d '{ "model": "kimi-k2.7-code", "messages": [{"role":"user","content":"Review this function"}] }'
The model value carries no provider prefix — the gateway resolves it to a provider (Workers AI, Moonshot's own API, or another host of the same weights).

  • Routing evolution ladder disclosed by the post:
  • Provider-first (today's default) — model string names the provider.
  • Model-first — name the model; gateway selects provider + fails over.
  • Smart routing — name nothing; Workers-AI classifier + heuristic scorer pick the model from a curated pool.

Operational numbers / status

  • Model-first routing: pilot "in the coming months" for all AI Gateway + Workers AI users.
  • Smart routing: currently piloting internally; testing/iterating "in the next few weeks before release."
  • Unified billing on Workers AI: shipped ("launching today"); elevated rate limits available on the unified-billing path (numbers in developer docs).
  • No latency/throughput/accuracy benchmarks for the classifier or scorer are given.

Caveats

  • Roadmap post, not a system-internals deep-dive. Model-first routing and smart routing are announced-as-coming; the classifier architecture ("a classifier running on Workers AI" + "a heuristic scorer") is described at a conceptual level with no model, feature set, or evaluation numbers.
  • Product-launch framing. Included per AGENTS.md borderline rule because the routing-architecture content (control-plane convergence, model-first vs provider-first routing, classifier-based smart routing, resiliency-as-routing) is well over 20% of the body and is genuinely architectural.
  • "Vetted providers" and ZDR handling are asserted, not detailed.

Source

Last updated · 766 distilled / 2,225 read