SYSTEM Cited by 11 sources
Workers AI¶
Overview¶
Workers AI is Cloudflare's managed LLM + embedding + reranker inference platform on the developer platform — models served from the @cf/… namespace, bound to a Worker via the ai binding:
Models include:
- LLMs —
@cf/moonshotai/kimi-k2.5(systems/kimi-k2-5) as the chat model in the 2026-04-16 AI Search support-agent example + the canonical extra-large-model-serving example in the 2026-04-16 high-performance-LLMs post. - Rerankers —
@cf/baai/bge-reranker-baseas the default cross-encoder reranker option for AI Search hybrid retrieval. - Embeddings, image generation, speech, … — the broader model catalog.
Role in the stack¶
- For the Agents SDK: invoked via
createWorkersAI({ binding: this.env.AI })+workersai("@cf/moonshotai/kimi-k2.5")passed to the VercelaiSDK'sstreamText. - For AI Search: the reranker stage runs a Workers AI model when
reranking: trueis set on an instance. Embeddings for the vector half are also Workers-AI-served (model not named in the 2026-04-16 post).
Billing: Workers AI and AI Gateway usage are billed separately from the AI Search product itself, even during AI Search's open beta.
High-performance serving stack (2026-04-16 post)¶
The "Building the foundation for running extra-large language models" post (2026-04-16) discloses the internal serving architecture behind Workers AI for extra-large models like Kimi K2.5 (>1T parameters, ~560 GB weights). Four load-bearing pieces:
1. Prefill/Decode disaggregation¶
Separate inference servers for the compute-bound prefill stage (processes input tokens, populates KV cache) and the memory-bound decode stage (generates output tokens). Non-trivial load balancer does:
- Two-hop routing (prefill → decode).
- KV-transfer metadata passing across stages.
- SSE response rewrite (decode's stream augmented with prefill-side cached-token counts before reaching the client).
- Token-aware admission — tracks in-flight tokens per stage pool separately.
Measured effect: p90 TTFT dropped; p90 intertoken latency: ~100 ms with high variance → 20-30 ms (3× improvement), using the same quantity of GPUs while request volume increased. See prefill-decode-disaggregation, disaggregated-inference-stages. (Source: sources/2026-04-16-cloudflare-building-the-foundation-for-running-extra-large-language-models)
2. Session-affinity prompt caching (x-session-affinity header)¶
Clients signal per-session opaque tokens; Workers AI routes continuation requests back to the replica with hot KV cache for the conversation prefix. Incentivised by discounted cached-token pricing. Integrated into agent harnesses via PRs like OpenCode #20744.
Measured effect: peak input-cache-hit ratio 60% → 80% after heavy internal users adopted the header. "A small difference in prompt caching from our users can sum to a factor of additional GPUs needed to run a model." See concepts/context-engineering, session-affinity-header. (Source: sources/2026-04-16-cloudflare-building-the-foundation-for-running-extra-large-language-models)
3. Cluster-wide shared KV cache over RDMA (Mooncake Transfer Engine + Mooncake Store + LMCache / SGLang HiCache)¶
For models that span multiple GPUs (required for extra-large models), KV cache is shared cluster-wide via RDMA over NVLink + NVMe-oF. Mooncake Store extends cache onto NVMe storage (longer session residency). LMCache or SGLang HiCache is the software layer that exposes the shared cache to the serving engine.
Consequence: "This eliminates the need for session aware routing within a cluster" — cross-cluster, x-session-affinity still matters. See rdma-kv-transfer, multi-gpu-serving, kv-aware-routing. (Source: sources/2026-04-16-cloudflare-building-the-foundation-for-running-extra-large-language-models)
4. Speculative decoding with NVIDIA EAGLE-3¶
Drafter model nvidia/Kimi-K2.5-Thinking-Eagle3 drafts N future tokens; Kimi K2.5 verifies them in one parallel forward pass. "In agentic use cases, speculative decoding really shines because of the volume of tool calls and structured outputs that models need to generate. A tool call is largely predictable — you know there will be a name, description, and it's wrapped in a JSON envelope." See concepts/speculative-decoding, systems/eagle-3. (Source: sources/2026-04-16-cloudflare-building-the-foundation-for-running-extra-large-language-models)
5. Proprietary Rust inference engine: Infire¶
Cloudflare's own engine (Rust, originally Birthday Week 2025) runs on both prefill and decode tiers. Multi-GPU via pipeline, tensor, and expert parallelism. Much lower activation-memory overhead than vLLM — can fit Llama 4 Scout on 2× H200 with >56 GiB KV room (~1.2M tokens), Kimi K2.5 on 8× H100 (not H200) with >30 GiB KV room. Sub-20s cold boot. Up to +20% tokens/sec vs baseline on unconstrained systems. "In both cases you would have trouble even booting vLLM in the first place." See systems/infire. (Source: sources/2026-04-16-cloudflare-building-the-foundation-for-running-extra-large-language-models)
Unified catalog + BYO model (2026-04-16 AI Platform post)¶
The 2026-04-16 "Cloudflare's AI Platform: an inference layer designed for agents" post repositions Workers AI from "catalog of @cf/… open-source models" to "one binding for any model from any provider, plus your own":
- Same
env.AI.run()binding calls third-party providers.env.AI.run('anthropic/claude-opus-4-6', { input: ... }, { gateway: { id: 'default' } })is a one-line change from@cf/...— see unified-inference-binding. Callers fall through AI Gateway transparently. - 70+ models across 12+ providers on the catalog — named: Alibaba Cloud, AssemblyAI, Bytedance, Google, InWorld, MiniMax, OpenAI, Pixverse, Recraft, Runway, Vidu — including image, video, and speech models (not just text LLMs). Canonical unified-model-catalog instance.
- REST API for non-Workers callers committed for the coming weeks.
- BYO-model via Replicate Cog containers. Customer writes
cog.yaml+predict.py:Predictor+cog build, pushes to Workers AI, and the model surfaces in the same AI Gateway catalog — see byo-model-via-container. Current scope: Enterprise + design-partner access; roadmap: customer-facing push APIs,wranglercommands, and faster cold starts via GPU snapshotting. - Strategic context: the Replicate team has joined the Cloudflare AI Platform team ("we don't even consider ourselves separate teams anymore"); Replicate models are being brought onto AI Gateway and the hosted models are being replatformed onto Cloudflare infrastructure — this is what explains the catalog expansion from text-LLM-dominated to multimodal.
- Network-topology argument: "When you call these Cloudflare-hosted models through AI Gateway, there's no extra hop over the public Internet since your code and inference run on the same global network." The
@cf/…path remains the fastest TTFT path for latency-critical agent workloads.
Agentic-workload tuning posture¶
The serving stack's tuning knobs are explicitly optimised for agentic traffic shape: large system prompt + tool descriptions + MCP server metadata + growing conversation history → input-heavy with long reusable prefixes. The two things that matter most:
- Fast input-token processing — PD disaggregation + cluster-wide KV sharing + session affinity address this.
- Fast tool-call generation — EAGLE-3 speculative decoding shines here because tool-call output is structurally predictable (JSON envelope, known schema).
"After our public model launch, our input/output patterns changed drastically again. We took the time to analyze our new usage patterns and then tuned our configuration to fit our customer's use cases." Kimi K2.5 made 3× faster post-launch by retuning configuration for observed traffic shape, not by more hardware. (Source: sources/2026-04-16-cloudflare-building-the-foundation-for-running-extra-large-language-models)
Weight-compression lever: Unweight (2026-04-17)¶
Cloudflare adds Unweight as the weight-side VRAM-reduction lever on Workers AI: Huffman coding on the redundant BF16 exponent byte paired with a custom reconstructive matmul kernel that feeds tensor cores directly from SMEM. Llama-3.1-8B results: ~22 % model-size reduction (~3 GB VRAM saved per instance); bit-exact lossless; H100-only at launch. Complements Infire's activation-memory discipline — Unweight on weights, Infire on activations, savings additive into KV-cache headroom + serving density. Open-source kernels ship as systems/unweight-kernels. (Source: sources/2026-04-17-cloudflare-unweight-how-we-compressed-an-llm-22-percent-without-sacrificing-quality)
Hosting Clef — decision models + non-autoregressive serving (2026-10-01)¶
Workers AI is the hosting + serving substrate for Clef and Clef-flash, Cloudflare's first in-house-trained ML models and its first decision models — typed, probabilistic classifiers rather than text generators. Two things make this a distinct Workers AI serving shape:
- Edge-GPU hot-path placement. Because Clef runs on "our GPUs at the edge," the decision model has low enough network latency to sit in the agent hot path — Cloudflare's framing is to put Clef in the hot path to decide and "combine that with one of our LLMs on Workers AI to take action" (the cheap decision model gating the expensive general model — patterns/cheap-approximator-with-expensive-fallback). Clef is the precision tier; Clef-flash is for latency-critical decisions.
- Non-autoregressive decision inference. Clef runs Qwen for a prefill-only pass, then scores valid schema choices in parallel — "the decision step is non-autoregressive, so there's no intermediate text to generate token by token," making it much faster than the autoregressive LLM serving the rest of the stack is tuned for. Measured: median 209.3 ms (Clef) / 38.8 ms (Clef-flash); and 2.2 s vs 4.7 s (gpt-oss-120b) on Cloudflare's threat-intel domain-classification workflow.
Workers AI is also the rollout + redeploy substrate for Clef's RL fine-tuning product: it generates rollouts against the base Clef model, and the fine-tuned model is redeployed via Workers AI + BYO Model (Cog) — see Clef for the full AI-Gateway → Workers-AI → Containers → Trainer loop. (Source: sources/2026-10-01-cloudflare-introducing-clef-our-open-source-decision-models-and-new-rl)
Seen in¶
-
sources/2026-10-01-cloudflare-introducing-clef-our-open-source-decision-models-and-new-rl — Workers AI hosts Cloudflare's first in-house-trained models, Clef + Clef-flash decision models. Edge-GPU hot-path placement for in-the-loop agent decisioning; non-autoregressive decision inference (prefill-only Qwen pass + parallel schema scoring) distinct from the autoregressive LLM serving stack; also the rollout-generation + BYO-Model redeploy substrate for Clef's RL fine-tuning loop.
-
sources/2026-09-28-cloudflare-emdash-10-the-stable-cms-with-a-secure-plugin-registry — Workers AI as the moderation engine for the EmDash plugin registry. EmDash 1.0's open-source labeler service "uses Workers AI to moderate package descriptions" — annotating the names/descriptions/links/images the catalog displays without rewriting or taking ownership of the underlying atproto publication. A content-moderation deployment of Workers AI at the registry layer.
-
sources/2026-05-19-cloudflare-announcing-claude-managed-agents-on-cloudflare —
image_generateagent tool delivered by Workers AI in the Claude Managed Agents integration. Quoting the post: "image_generate, which uses Workers AI to generate images on Cloudflare. This pairs well with Claude providing text-based inference." Architectural shape: the brain (Claude) handles text reasoning and tool selection; image generation is offloaded to a Cloudflare-hosted multimodal model, called via the operator's outbound Worker proxy. Multimodal-via-tool-call rather than multimodal-via- single-model. - sources/2026-04-16-cloudflare-ai-platform-an-inference-layer-designed-for-agents — canonical unified-catalog + BYO-model launch. Same
env.AI.run()binding previously scoped to@cf/…now calls 70+ models across 12+ providers with a one-line provider swap; BYO-model via Cog containers; automatic provider failover and buffered resumable streaming become gateway-owned concerns so the caller doesn't write retry logic. Multimodal catalog expansion (image, video, speech) alongside text LLMs. - sources/2026-04-16-cloudflare-ai-search-the-search-primitive-for-your-agents — named as the inference substrate for both the chat LLM (Kimi K2.5) and the hybrid-search cross-encoder reranker (bge-reranker-base).
- sources/2026-04-16-cloudflare-building-the-foundation-for-running-extra-large-language-models — deep dive on the high-performance serving stack: PD disaggregation, session affinity, Mooncake-based cluster KV sharing, EAGLE-3 speculative decoding, Infire proprietary inference engine, multi-GPU parallelism for extra-large models. Canonical wiki instance of the five serving primitives above.
- sources/2026-04-16-cloudflare-email-service-public-beta-ready-for-agents — Workers AI as the inbound email classification substrate for Agentic Inbox. The reference app runs inbound mail through a Workers-AI classifier before persisting + replying — canonical wiki instance of the "classify" stage of inbound-classify-persist-reply-pipeline.
- sources/2026-04-17-cloudflare-unweight-how-we-compressed-an-llm-22-percent-without-sacrificing-quality — canonical lossless-weight-compression lever on Workers AI; ~22 % model-size reduction on Llama-3.1-8B via Huffman coding on BF16 exponents + fused tensor-core reconstruction. H100 only at launch. Pairs with Infire's activation-memory discipline.
Related¶
- systems/cloudflare-workers — host runtime + binding layer.
- systems/cloudflare-ai-gateway — optional proxy / observability layer in front of Workers AI or third-party providers. As of 2026-04-16 the unifying catalog surface for first-party + third-party + BYO models.
- systems/cloudflare-ai-search — consumer of the reranker + embedding models.
- systems/replicate-cog — BYO-model container format for pushing custom inference code to Workers AI.
- systems/kimi-k2-5 — specific extra-large model served.
- systems/cloudflare-agents-sdk — canonical consumer in agent workloads.
- systems/infire — Cloudflare's Rust inference engine inside Workers AI.
- systems/unweight — lossless weight-compression lever complementing Infire's activation-memory discipline.
- systems/unweight-kernels — open-source CUDA kernels backing Unweight.
- systems/mooncake-transfer-engine / systems/mooncake-store — KV transport + NVMe tier.
- systems/eagle-3 — speculative-decoding drafter.
- systems/lmcache / systems/sglang — cluster-wide cache layers.
- unified-model-catalog — the product-surface property the 2026-04-16 launch realises on Workers AI.
- concepts/retrieval-ranking-funnel — reranker model served on this platform.
- prefill-decode-disaggregation / concepts/context-engineering / concepts/speculative-decoding / multi-gpu-serving / rdma-kv-transfer / concepts/tensor-parallelism / concepts/pipeline-parallelism / concepts/mixture-of-experts
- unified-inference-binding — the
env.AI.run()one-line-swap pattern. - byo-model-via-container — the Cog-based BYO substrate pattern.
-
companies/cloudflare — parent org.
-
sources/2026-08-04-cloudflare-astro-issue-triage — Astro issue-triage factory model endpoint. The publicly shown Action configuration uses Workers AI model identifiers
@cf/moonshotai/kimi-k2.7-codefor triage and@cf/moonshotai/kimi-k2.6for verification. The post discloses configuration names only, not performance, cost, model-routing policy, or the security boundary around issue text. (Source: sources/2026-08-04-cloudflare-astro-issue-triage)
Unification with AI Gateway into one control plane (2026-08-07)¶
Workers AI and AI Gateway are being unified
into a single AI control plane: the two products share the env.AI.run()
binding and a unified /ai/ REST API, so a Workers AI call becomes a gateway
call by adding { gateway: { id: 'default' } }. Three consequences for Workers
AI specifically:
- Observability by default. Routing Workers AI calls through the
defaultgateway yields per-request logging, per-model token counts, and cost attribution with no setup (concepts/centralized-ai-governance). - Unified billing + elevated rate limits. Workers AI usage can now be paid from a shared prepaid AI Gateway wallet alongside external providers (patterns/unified-billing-across-providers); the unified-billing path also raises Workers AI rate limits.
- First-party host in model-first / smart routing. Under model-first routing, Workers AI is the preferred host when it has capacity for a named model, with transparent spillover to other vetted providers on saturation (model-first-routing-with-transparent-failover). Workers AI also hosts the classifier that powers smart routing (predicting task type/complexity/context before a model is chosen).
(Source: sources/2026-08-07-cloudflare-unifying-workers-ai-and-ai-gateway)
Hosting the Auto Router classifier (2026-09-30)¶
Workers AI is the substrate for AI Gateway's
Auto Router (cloudflare/auto). The routing decision is made by a multi-head
classification model running on Workers AI and deployed on GPUs across Cloudflare's
edge network — not by the frontier model being routed to. Per request, the classifier
reads a compact, newest-turns-weighted view of the conversation and emits (1)
probabilities over 14 task categories and (2) 1–5 ratings across four dimensions
(complexity, ambiguity, stakes, dependence on earlier context); a downstream scoring
matrix + utility function then selects the serving model. This makes Workers AI the
cheap, edge-local first-pass model that decides how to spend on the expensive one —
a routing-classifier role distinct from Workers AI's usual role as the served model.
(Source: sources/2026-09-30-cloudflare-cut-your-ai-spend-with-ai-gateways-auto-router)
Seen in¶
- sources/2026-09-30-cloudflare-cut-your-ai-spend-with-ai-gateways-auto-router —
Workers AI hosts the multi-head request classifier behind AI Gateway's
cloudflare/autoAuto Router (14 task categories + 4 difficulty dimensions, on edge GPUs). - sources/2026-10-01-cloudflare-ai-search-is-now-generally-available —
Workers AI hosts
Qwen3-VL-Embedding, the native-multimodal embedding model behind AI Search's GA image retrieval; it embeds image pixels and text into one vector space. Embedding + reranking on default Workers AI models are free inside AI Search's included pipeline. The richer multimodal embeddings are kept cheap with MRL truncation.