Skip to content

CLOUDFLARE 2026-10-01 Tier 1

Read original ↗

Introducing Clef: our open-source decision models, and new RL fine-tuning platform

Summary

Cloudflare released Clef and Clef-flash, two Cloudflare-trained decision models — models that produce bounded, strictly-typed, probabilistic structured outputs cheaply and consistently when a workflow needs a classification/decision, as opposed to the open-ended, non- deterministic text generation of a general LLM. They are hosted on Workers AI, fully Jev-API-compatible, open-sourced on Hugging Face under Apache 2.0, and (per Cloudflare) lead the Jev Decision Index. The post also debuts a reinforcement-learning (RL) fine-tuning product that lets customers tune Clef to a workload — first hands-on via a forward-deployed-engineer (FDE) team, later as a self-serve platform — assembled entirely from existing Cloudflare AI-platform primitives (AI Gateway dataset capture → Workers AI rollouts → Containers RL sandbox → a new Trainer → redeploy via Workers AI + BYO Model (Cog)). The serving-infra and training-infra depth — non-autoregressive decision inference, edge-GPU hot-path placement, frozen-Qwen-backbone + rank-256 LoRA post-training, and the RL-platform composition — is what puts this launch post decisively in scope.

Key takeaways

  1. A decision model is a typed, probabilistic classifier for agent control flow. You pass inputs (e.g. a support message) + a schema of questions (noul boolean, choice, score) and get back typed answers with probabilities your code uses to route, escalate, or defer to a human — so "a human does not necessarily need to be in the loop for agentic decisions anymore" while still being able to defer when needed (concepts/human-in-the-loop, concepts/structured-output-reliability). (Source: this article)

  2. Non-autoregressive decision inference is the latency win. During inference Clef runs Qwen for a prefill-only pass, then scores the valid schema choices in parallel — "the decision step is non-autoregressive, so there's no intermediate text to generate token by token." Rather than generating text, Clef/Clef-flash derive schema choices directly from internal backbone representations via a specialized two-stage attention routing process (every valid choice extracts prompt-relevant context; field parameters cross-attend with other fields and back to the payload before scoring; a lexical prior preserves semantic intent). The architecture "unites option-specific evidence routing, joint cross-field attention, and schema-bound scoring." (Source: this article)

  3. Edge-GPU hosting puts a decision model in the agent hot path. Because Clef is hosted on Workers AI, it runs on Cloudflare's GPUs at the edge → low network latency → "you could put Clef into the hot path for agents to make decisions and combine that with one of our LLMs on Workers AI to take action." Clef is the precision model; Clef-flash is for latency-critical decisions. (Source: this article)

  4. Measured production win on Cloudflare's own Threat Intelligence workload. Classifying a domain's categories (with Browser Run) took Clef 2.2 s to fetch + render + classify, vs 4.7 s for the fastest general LLM gpt-oss-120b in the same workflow — and gpt-oss returned only two classifications. (Source: this article)

  5. Latency benchmarks (43 evals): Clef median 209.3 ms / p95 238.6 ms; Clef-flash median 38.8 ms / p95 122.4 ms; vs Jev median 524.1 ms. (Laya is faster — median 5.8 ms — but trades off quality.) (Source: this article)

  6. Trained by post-training a frozen Qwen backbone with LoRA + calibration losses. "By freezing Qwen3.8-27B for Clef and Qwen3.5-9B for Clef-flash, we jointly optimized the routing head alongside rank-256 low-rank adapters" (LoRA). Post-training used label-smoothed cross-entropy for valid schema outputs paired with a Brier loss to refine probability calibration, over internal synthetic datasets that permutate field orders, prompts, and schema structures. (Source: this article)

  7. RLCD = Reinforcement Learning for Calibrated Decisions. A secondary optimization target that grants partial credit to adjacent ordinal choices, rewards fully-precise record outputs, and applies a reference penalty to prevent distribution shift — for better accuracy and generalization. (Source: this article)

  8. The RL product is a composition of existing AI-platform primitives, not new infrastructure. "We've been building our AI platform to have the right primitives where we could be building a custom RL product." The loop: AI Gateway proxies all AI traffic and automatically creates a dataset of requests → Workers AI generates rollouts against the base Clef model → Containers provide the RL sandbox for scoring and replaying agent actions → a [NEW] Trainer updates fine-tuned-Clef weights → redeploy via Workers AI + BYO Model (Cog, the Replicate acquisition). (Source: this article)

  9. Fine-tuning on 15+ years of labelled network data is the internal driver. Internal use cases — Trust & Safety submission evaluation, Support-request triage, good-bot/bad-bot crawler decisions in the Bot products — want classifiers trained on Cloudflare's years of labelled decisions; fine-tuning trades general-purpose performance for higher domain accuracy. These are the first remit of the FDE fine-tuning team and the basis for the RL product. (Source: this article)

  10. Differentiators vs other decision models: a vision encoder (image classification; Jev is text-only), a 64k context window (Jev 32k), and enterprise privacy ("we don't read, store, or train on your requests or responses" unless you opt into fine-tuning). (Source: this article)

Systems / concepts / patterns extracted

  • New system (proper noun): Clef — Cloudflare's decision-model family (Clef + Clef-flash) + the RL fine-tuning platform.
  • Enriched systems: Workers AI (edge-GPU decision- model serving + non-autoregressive decision inference + Clef/Clef-flash hosting + RL rollouts), AI Gateway (traffic → RL training dataset capture), Containers (RL sandbox for scoring/replaying agent actions), Replicate Cog (BYO-Model redeploy of fine-tuned Clef), Qwen (Clef backbone — Qwen3.8-27B / Qwen3.5-9B frozen).
  • Mapped INTO existing concepts (taxonomy gate — no new pages): concepts/structured-output-reliability (bounded typed probabilistic schema outputs), concepts/lora-low-rank-adaptation (rank-256 adapters + frozen backbone), concepts/knowledge-distillation (post-training a general backbone into a specialized classifier), concepts/hot-path (decision model in the per-request agent path), concepts/latency-critical-vs-latency-tolerant-workload (Clef precision vs Clef-flash latency-critical), concepts/human-in-the-loop (programmatic decide/escalate/defer), concepts/model-first-routing (decision model as the cheap classifier that gates the expensive LLM).
  • Mapped INTO existing patterns: patterns/cheap-approximator-with-expensive-fallback (cheap edge decision model gating an expensive general LLM "to take action"), patterns/teacher-student-model-compression (specialize a frozen general backbone into a small fast classifier).
  • Taxonomy gate — did NOT mint pages for non-autoregressive decision inference, two-stage attention routing, schema-bound scoring, RLCD (Reinforcement Learning for Calibrated Decisions), Brier-loss probability calibration, or decision model as a term. These are Clef-specific training/inference mechanics (single source here) and are recorded as prose + tags on the Clef page and in this source; candidates for Lint promotion if a second distinct source corroborates. Trainer (the new weight-update component) recorded as prose on the Clef page (unnamed internal product, single-source).

Operational numbers

  • Latency: Clef median 209.3 ms / p95 238.6 ms; Clef-flash median 38.8 ms / p95 122.4 ms; Jev median 524.1 ms / p95 536.0 ms.
  • Threat-intel workflow: Clef 2.2 s vs gpt-oss-120b 4.7 s (fetch + render + classify a domain) — ~2× faster, more classifications returned.
  • Backbones: Clef = Qwen3.8-27B (frozen); Clef-flash = Qwen3.5-9B (frozen); rank-256 LoRA adapters + a routing head jointly optimized.
  • Context window: 64k (vs Jev's 32k).
  • Benchmarks: 43 eval benchmarks run; Clef leads several (e.g. BANKING77 macro-F1 94.20, CLINC150+OOS 97.43, ToolRet nDCG@10 69.19); beats Jev on 3/4 of Typesafe's workflow evals.
  • Example domain classification: 95% fashion, 85% ecommerce, <1% phishing.
  • License: Apache 2.0, weights on Hugging Face (Cloudflare/clef, Cloudflare/clef-flash).

Caveats

  • Launch post: benchmark results are Cloudflare-reported on the Jev Decision Index + Typesafe's eval suite; no independent third-party reproduction.
  • The RL fine-tuning product is FDE-hands-on first; self-serve is roadmap ("later as a self-serve fine-tuning platform"). Trainer is named but not architecturally detailed.
  • Two-stage attention routing / schema-bound scoring are described narratively; no reference implementation or paper is linked for the routing architecture itself (weights are open, mechanism prose is high-level).
  • Qwen3.8-27B / Qwen3.5-9B are the post's version strings — treat as the article's naming, not necessarily canonical Alibaba release labels.

Source

Last updated · 766 distilled / 2,225 read