Skip to content

SYSTEM Cited by 10 sources

Cloudflare AI Gateway

AI Gateway is Cloudflare's LLM proxy tier that sits between an application and any of the supported AI providers (Anthropic, OpenAI, Google, Workers AI, and others). It gives developers centralised visibility, per-request logging, cost tracking, rate limiting, and — configured as a drop-in — the ability to change the LLM model or provider without redeploying application code (see patterns/ai-gateway-provider-abstraction).

Core capabilities

  • Unified catalog + unified binding (2026-04-16). All models — first-party Workers AI @cf/…, 70+ third-party models across 12+ providers (Anthropic, OpenAI, Google, Alibaba Cloud, AssemblyAI, Bytedance, InWorld, MiniMax, Pixverse, Recraft, Runway, Vidu), and BYO Cog containers — are callable through one env.AI.run(model_string, ...) binding. Provider selector lives inside the model string; switching providers is a one-line edit. REST API for non-Workers callers promised in the weeks after launch (unified-model-catalog, unified-inference-binding).
  • Multimodal, not just text. Image, video, and speech models surface alongside text LLMs in the same catalog.
  • Provider abstraction. Application points at the gateway endpoint (e.g. ANTHROPIC_BASE_URL=<gateway>); swapping models/providers happens in gateway config, not application deploys.
  • Bring-Your-Own-Key (BYOK). Instead of shipping provider secrets with every request, Secrets Store integration lets Cloudflare inject the key server-side on behalf of the caller (concepts/byok-bring-your-own-key).
  • Unified Billing. Alternative to BYOK: top up a Cloudflare account with credits and Cloudflare pays providers directly, deducting from the credit balance — no provider-secrets management at all. Per-request custom metadata (metadata: { teamId, userId, ... }) enables spend attribution by user / tenant / workflow.
  • Automatic provider failover (2026-04-16). When a model is available on multiple providers and one goes down, the gateway silently routes to another — the application never writes retry-on-outage logic (automatic-provider-failover). Extends the configured fallback chain feature into an automatic cross-provider retry primitive.
  • Buffered resumable streaming (2026-04-16). Streaming inference responses are buffered gateway-side, independently of the caller's lifetime. If an Agents SDK agent is interrupted mid-inference, it reconnects and retrieves the buffered response — no re-inference, no double-billing. Paired with Agents SDK checkpointing, "the end user never notices" a mid-turn crash (resilient-inference-stream, buffered-resumable-inference-stream).
  • BYO-model via Cog containers (2026-04-16). Customers build Replicate Cog containers (cog.yaml + predict.py:Predictor) and push them to Workers AI; the gateway surfaces them alongside first-party and third-party models in the same catalog. Currently Enterprise + design-partner access (byo-model-via-container).
  • Colo-with-inference latency. Cloudflare's 330-city network means the gateway sits close to both users and inference endpoints; for @cf/… models the entire call path stays inside Cloudflare (no public-Internet hop), which is the load-bearing first-token-latency argument for agent workloads (concepts/tail-latency-at-scale).
  • Provider / model fallbacks. Ordered list of providers/models tried on failure; the application's retry logic becomes a gateway config change.
  • Observability. Request/response logs, spend by model, latency distribution, error rates — the same view across heterogeneous upstream providers.

Seen in

  • sources/2026-04-16-cloudflare-ai-platform-an-inference-layer-designed-for-agents — canonical unified-catalog launch. Same env.AI.run(...) binding previously scoped to Workers AI @cf/… models now calls any of 70+ models across 12+ providers — one line of code to switch, one set of credits to pay. Post framing: "Today, we're making Cloudflare into a unified inference layer: one API to access any AI model from any provider, built to be fast and reliable." Same post introduces automatic provider failover (gateway silently routes to a second provider when the first goes down — automatic-provider-failover), buffered resumable streaming (stream survives agent disconnects — resilient-inference-stream, buffered-resumable-inference-stream), and BYO-model via Cog containers (byo-model-via-container) pushed to Workers AI. Strategic context: Replicate team joined the Cloudflare AI Platform team, bringing multimodal-model-hosting DNA (image, video, speech) that explains the catalog expansion from text-LLM-dominated to multimodal. Named new providers: Alibaba Cloud, AssemblyAI, Bytedance, Google, InWorld, MiniMax, OpenAI, Pixverse, Recraft, Runway, Vidu.
  • sources/2026-01-29-cloudflare-moltworker-self-hosted-ai-agent — canonical example of the zero-code-change provider swap. Porting Moltbot to run on Workers required only setting ANTHROPIC_BASE_URL to the AI Gateway endpoint; Moltbot's code was unchanged, and the upgrade path to BYOK or Unified Billing is a gateway-config operation.
  • sources/2026-04-20-cloudflare-internal-ai-engineering-stack — AI Gateway as Cloudflare's internal platform-layer choke point: every LLM call from internal tooling (OpenCode, Windsurf, AI Code Reviewer, Dynamic Workers) flows through a Worker that validates a Zero Trust Access JWT, tags the request with an anonymous per-user UUID, and forwards to AI Gateway. Reported: 20.18M req/month, 241.37B tokens, 91% frontier labs / 9% Workers AI.
  • sources/2026-04-16-cloudflare-ai-search-the-search-primitive-for-your-agents — AI Gateway is called out as a still-separately-billed companion during AI Search's open beta alongside Workers AI: "Workers AI and AI Gateway usage will continue to be billed separately." Sits alongside AI Search as the LLM-proxy tier in the agent stack (retrieval is AI Search; inference is Workers AI through AI Gateway).

Standards-review routing

Cloudflare's spec reviewer routes model requests through AI Gateway while it evaluates filtered Codex requirements against design documents. This is a concrete internal workflow-review consumer alongside AI Code Review: the gateway carries inference for a scheduled, stateful Worker/D1 review service rather than only interactive developer tooling. The source does not disclose model selection, gateway policy, or traffic volume. (Source: sources/2026-08-04-cloudflare-how-cloudflare-enforces-engineering-standards-using-ai)

Seen in

Identity-aware gateway analytics (2026-08-05)

AI Gateway can now sit behind Cloudflare Access on a custom domain. Access authenticates the caller through a SAML-capable identity provider, and the gateway carries the verified Access user ID in cf.user_id metadata on every request. This makes the central-proxy boundary simultaneously a per-user access, spending, logging, and behavioral-analytics boundary rather than merely a provider router. User Insights builds a rolling 30-day per-account p95 session-cost baseline from that traffic and alerts only when a session also crosses the organisation-level p99 cost floor; it presents candidates for human review rather than blocking the account. (Source: sources/2026-08-05-cloudflare-catching-rogue-ai-behavior-with-identity-aware-analytics)

The source does not disclose the sessionization method, history required before a baseline is eligible, alert precision/recall, default thresholds, or data-retention policy. The published $200 p99 is an internal illustration, not a service default.

Cloudflare OS inference-governance tier (2026-08-05)

Cloudflare OS routes every inference request through AI Gateway. The published role is organizational governance: administrators choose available models and model selection per job, attribute requests to a person/team/workspace, set budgets and rate limits, and decide behavior when a limit is reached. This adds an agent-workspace cost-control deployment alongside the existing provider-routing, internal engineering, and identity-aware analytics uses. The article does not disclose limits, model-routing algorithms, or over-limit behavior. (Source: sources/2026-08-05-cloudflare-os-an-open-platform-for-agents-apps-and-work)

Unification with Workers AI into one control plane (2026-08-07)

AI Gateway and Workers AI converge into a single AI control plane: there is no separate binding for each — both flow through the same env.AI.run(model, input, options) path and a unified /ai/ REST API, so choosing between "gateway" and "first-party inference" is no longer a first-order decision. Three concrete pieces:

  • default gateway → zero-setup observability. Passing { gateway: { id: 'default' } } (or cf-aig-gateway-id: default) auto-creates a gateway on the first authenticated request; every request is then logged with full payloads, per-model token counts, and cost attribution without any dashboard setup. This makes centralized AI governance the opt-out default for all Workers AI traffic. Outgrow it by creating a named gateway and changing one parameter.
  • Unified billing now covers Workers AI. A single prepaid AI Gateway wallet is spendable across OpenAI, Anthropic, Workers AI, or any supported provider (patterns/unified-billing-across-providers); the unified-billing path also grants elevated Workers AI rate limits (numbers in developer docs).
  • Model-first + smart routing (roadmap). Model-first routing lets a caller name a model with no provider prefix ("model": "kimi-k2.7-code") and have the gateway resolve the provider, load-balance across hosts of the same weights, and fail over transparently (model-first-routing-with-transparent-failover) — piloting in the coming months. On top of it, smart routing runs a classifier on Workers AI to predict task/complexity/context and a heuristic scorer to pick the model from a curated pool with zero config (patterns/specialized-agent-decomposition) — piloting internally. Both reframe resiliency as a routing problem handled at the central choke point rather than by per-Worker retry logic. (Source: sources/2026-08-07-cloudflare-unifying-workers-ai-and-ai-gateway)

Seen in

User Insights: model-fit context + async classification (2026-09-30)

Alongside the Auto Router, User Insights gains a model-fit view over already-proxied traffic: a model overkill view (flags conversations where the chosen model is more capable/costlier than the task needs, attributed to specific users/agents/apps), task analysis (groups conversations into coding / research / writing / summarization / data-analysis), turns analysis (back-and-forth count + cumulative time/tokens/cost per session), and a Potential Savings view (requests a faster/cheaper model could likely handle). It is a starting point for investigation — it does not auto-recommend a replacement or block traffic. Free to AI Gateway users. (Source: sources/2026-09-30-cloudflare-identify-ai-model-overuse-with-user-insights)

The signals come from a dedicated categorization Worker running asynchronously, off the request path: AI Gateway writes the log to its existing storage path first, then the Worker classifies the conversation trajectory (user/assistant turns, tool calls, tool results) into a task category + confidence + four scored dimensions (complexity, intent ambiguity, stakes, context dependence), and joins the result to the log metadata. Because classification is out-of-band it adds no latency to the user's response, at the cost of the dashboard not being real-time (analysis trails traffic by ~1 day) — a deliberately latency-tolerant placement. It reuses AI Gateway's existing log architecture: metadata separate from log bodies, currently Durable Objects for metadata and R2 for log bodies. These same signals feed the Auto Router below — User Insights is the observe half, the Auto Router the act half. The post does not disclose which model powers the Worker, its accuracy, or log-eligibility rules; the storage split and ~1-day lag are stated as current-implementation details. (Source: sources/2026-09-30-cloudflare-identify-ai-model-overuse-with-user-insights)

Auto Router — cloudflare/auto (2026-09-30, public beta)

Auto Router productizes the previously-roadmapped smart routing into a shippable feature: set the model to cloudflare/auto and AI Gateway picks a model "capable enough for the task" per request, with no user model selection. It is the concrete realization of model-first routing as zero-config model selection. Internal use (OpenCode harness + Cloudflare OS) reports up to 30% cost savings vs. frontier-only. Free during beta.

How it works (two-stage, legible-by-construction):

  1. Pool construction (filter chain). The gateway builds the set of models that can serve the request: filters on request format / execution mode, then applies the credentials, billing config, access-control policies, and spend limits attached to the gateway, and drops unhealthy upstreams (auto-restored after outage). This is the same pool/eligibility machinery behind provider abstraction and automatic failover.
  2. Multi-head classifier on Workers AI (edge GPUs). Reads a compact view of the conversation (newest turns weighted) and emits (a) probabilities over 14 task categories (coding, planning, research, data analysis, …) and (b) 1–5 ratings on four dimensions: complexity, ambiguity, stakes, dependence on earlier context.
  3. Scoring matrix. Combines classifier signals with per-model benchmark results to estimate each candidate's fit. Calibrated by hand-defining the preferred model for example task/difficulty profiles and tuning weights to reproduce those choices. Adding a new model = adding its benchmark-derived weights; no classifier retraining.
  4. Utility selection. utility = expected quality − adaptive cost penalty. On easy requests the price term dominates (a smaller model can win); as difficulty rises the cost penalty falls and stronger models win. Returns a ranked list; the gateway tries the winner and can fall through to another eligible model.

Cache-aware model switching (the novel cost model). For long agentic sessions, cost is dominated by KV-cache read cost, which grows with context length. Switching models discards the hot cache and forces a full context re-write, so Auto Router prices switching explicitly:

  • Within a turn the cache is hot → switching rarely pays off → keep the same model.
  • Across turns a switching penalty grows with tokens already in context. A model holding a live session cache is priced at its cheap cache-read rate; every other candidate is priced at the full context-rewrite cost. "The deeper the conversation, the more a switch has to earn back."
  • Reasoning tokens compound it: most models can't read another model's reasoning tokens, so a switch that drops them forces redo at output prices. Roadmap: prefer the same model family on a switch.

Framings worth citing: the "jagged frontier" (the solution exists somewhere in the model portfolio; the router picks the right one per task, concepts/efficiency-frontier) and "lower token prices do not always produce lower-cost outcomes" — the router minimizes predicted trajectory cost, not dollars-per-million-tokens load balance.

Multiple routing profiles from one classifier: planned cloudflare/auto-best reuses the same classification + pool but selects highest expected quality with no cost tradeoff. Roadmap also adds zero-data-retention filtering, provider-capacity awareness, per-request reasoning-level selection, and Responses API / WebSockets support. (Source: sources/2026-09-30-cloudflare-cut-your-ai-spend-with-ai-gateways-auto-router)

Seen in

AI Gateway as the RL training-dataset capture layer (2026-10-01)

In Cloudflare's Clef RL fine-tuning product, AI Gateway is the first stage of the RL loop: "pass all your AI traffic through AI Gateway and automatically create a dataset of requests for your use case." This is a new role for the central-proxy choke point — beyond observe / govern / route (provider abstraction, cost control, Auto Router), the proxied traffic log becomes the training corpus for fine-tuning a decision model to a customer's workload. The downstream loop is Workers AI (rollouts) → Containers (RL sandbox) → Trainer (weight update) → Workers AI + BYO Model (Cog) redeploy. The choke-point-as-dataset-source shape is the same corpus-mining posture AI Gateway User Insights uses for out-of-band model-fit analysis — here applied to training rather than analytics. (Source: sources/2026-10-01-cloudflare-introducing-clef-our-open-source-decision-models-and-new-rl)

Seen in

Last updated · 766 distilled / 2,225 read