Skip to content

SYSTEM Cited by 1 source

Clef

Clef (and its smaller sibling Clef-flash) is Cloudflare's family of open-source decision models — models that produce bounded, strictly- typed, probabilistic structured outputs cheaply, quickly, and consistently when a workflow needs a classification or decision, in contrast to the open- ended, non-deterministic text generation of a general-purpose LLM. Launched 2026-10-01, hosted on Workers AI, fully Jev-API- compatible, and open-sourced on Hugging Face (Cloudflare/clef, Cloudflare/clef-flash) under Apache 2.0. Cloudflare's first in-house- trained ML model from the Workers AI team.

The name is a music-theory metaphor: a clef assigns pitch to the lines of a staff — it "helps define the domain of the context and the subsequent notes (actions) that follow it," and the CF hearkens to Cloudflare.

What a decision model is

Pass inputs (e.g. a customer-support message as state) plus a schema of questions, and get back typed answers with probabilities your code acts on. The API (@cf/cloudflare/clef) supports question types including:

  • noul — a boolean (e.g. "is this support request urgent?"),
  • choice — pick one of named criteria (e.g. billing / technical / sales),
  • score — an ordinal rating over a criteria ladder (e.g. No impact → Critical).

Because outputs are typed + probabilistic, agents can programmatically gather context, decide, and act — or defer to a human when needed, removing the human from the loop for routine agentic decisions (concepts/human-in-the-loop, concepts/structured-output-reliability).

How Clef is different

  • Vision encoder — classifies image/visual content (Jev is text-only).
  • 64k context window (vs Jev's 32k) — more input state to classify against.
  • Accuracy — competitive-to-leading across the Jev Decision Index and Typesafe's workflow evals (Cloudflare-reported; beats Jev on 3/4 workflow areas; Clef-flash notably strong given its speed).
  • Enterprise privacy — Cloudflare "don't read, store, or train on your requests or responses" unless you opt into the fine-tuning product.
  • Two tiers — Clef is the precision model; Clef-flash is for latency-critical decisions.

Inference architecture — non-autoregressive decision scoring

The latency win is architectural, not just model-size. During inference Clef:

  1. Runs Qwen for a prefill-only pass over the inputs.
  2. Scores the valid schema choices in parallel.

The decision step is non-autoregressive — "there's no intermediate text to generate token by token," which makes Clef significantly faster than autoregressive LLMs. Rather than generating text to produce structured answers, Clef and Clef-flash derive schema choices directly from internal backbone representations via a specialized two-stage attention routing process:

  • every valid choice extracts context relevant to the prompt;
  • individual field parameters cross-attend with other fields and back to the original payload before scoring;
  • a lexical prior preserves semantic intent across options.

The architecture "unites option-specific evidence routing, joint cross-field attention, and schema-bound scoring." (These are Clef-specific mechanics, recorded here as prose rather than minted as separate concept pages — see the source's taxonomy gate.)

Serving on Workers AI — edge GPU + hot path

Clef is hosted on Workers AI, so it runs on Cloudflare's GPUs at the edge → low network latency → "you could put Clef into the hot path for agents to make decisions and combine that with one of our LLMs on Workers AI to take action." This is the canonical shape of a cheap, fast decision model gating an expensive general LLM (patterns/cheap-approximator-with-expensive-fallback; a model-first-routing-adjacent split where the classifier decides whether/how to invoke the big model).

Measured: on Cloudflare's own Threat Intelligence domain-classification workflow (with Browser Run), Clef took 2.2 s to fetch + render + classify a website vs 4.7 s for the fastest general LLM gpt-oss-120b in the same workflow — ~2× faster and returning more classifications. Example output: a domain at 95% fashion, 85% ecommerce, <1% phishing.

Latency (43 evals): Clef median 209.3 ms / p95 238.6 ms; Clef-flash median 38.8 ms / p95 122.4 ms; vs Jev median 524.1 ms. (Laya is faster at median 5.8 ms but trades off quality.)

How Clef was trained

  • Frozen backbone + adapters. "By freezing Qwen3.8-27B for Clef and Qwen3.5-9B for Clef-flash, we jointly optimized the routing head alongside rank-256 low-rank adapters" (LoRA). The general Qwen base is post-trained / specialized into a decision classifier (concepts/knowledge-distillation-adjacent; patterns/teacher-student-model-compression shape).
  • Calibration losses. Label-smoothed cross-entropy for valid schema outputs paired with a Brier loss to refine probability calibration — the probabilities Clef returns are meant to be trustworthy, not just argmax-correct.
  • Synthetic data. Internal synthetic datasets that permutate field orders, prompts, and schema structures (robustness to schema shape).
  • RLCD — Reinforcement Learning for Calibrated Decisions. A secondary optimization target that grants partial credit to adjacent ordinal choices, rewards fully-precise record outputs, and applies a reference penalty to prevent distribution shift, improving accuracy + generalization.
  • Origin. Builds on a Cloudflare demo adapting DiffusionGemma to output deterministic probabilities by exposing LLM logprobs (itself building on Matt Mastracci's independent vLLM DiffusionGemma work); Clef switched the backbone to Qwen.

The RL fine-tuning product

Cloudflare offers a service to help customers fine-tune Clef to a workload — first hands-on with the forward-deployed-engineer (FDE) team, then (roadmap) as a self-serve platform customers use to capture data, fine- tune, and redeploy on Cloudflare. The internal driver: Cloudflare's 15+ years of labelled network data across domains lets it fine-tune far more accurate + faster specialized classifiers than generic Clef (Trust & Safety submission evaluation, Support triage, good-bot/bad-bot crawler decisions).

Crucially, the RL product is assembled from existing AI-platform primitives rather than new infrastructure — "we've been building our AI platform to have the right primitives where we could be building a custom RL product." The loop:

  1. Cloudflare AI Gateway — proxy all AI traffic through it and automatically create a dataset of requests for your use case.
  2. Cloudflare Workers AI — generate rollouts against the base Clef model.
  3. Cloudflare Containers — the RL sandbox for scoring and replaying agent actions.
  4. [NEW] Trainer — update the weights of the fine-tuned Clef model. (Named but not architecturally detailed in the launch post.)
  5. Workers AI + BYO Model (Cog) — redeploy the fine-tuned model on Workers AI (the Bring-Your-Own-Model work progressing since the Replicate acquisition).

Caveats

  • Benchmarks are Cloudflare-reported (Jev Decision Index + Typesafe eval suite); no independent reproduction.
  • The two-stage attention routing / schema-bound scoring are described narratively; weights are open but the routing-architecture mechanism prose is high-level, with no linked paper for the routing itself.
  • RL product is FDE-hands-on first; self-serve + Trainer internals are roadmap.
  • Qwen3.8-27B / Qwen3.5-9B are the post's version strings — the article's naming, not necessarily canonical Alibaba release labels.

Seen in

  • sources/2026-10-01-cloudflare-introducing-clef-our-open-source-decision-models-and-new-rl — canonical wiki introduction. Clef + Clef-flash decision models (Apache 2.0, Jev-API-compatible, hosted on Workers AI edge GPUs); non-autoregressive decision inference (prefill-only Qwen pass + parallel schema scoring + two- stage attention routing + schema-bound scoring); frozen-Qwen + rank-256 LoRA
  • label-smoothed-CE + Brier-loss calibration + RLCD training; the AI-Gateway → Workers-AI → Containers → Trainer → BYO-Model RL fine-tuning loop; 2.2 s vs 4.7 s threat-intel win; median 209 ms / Clef-flash 38.8 ms latency.
Last updated · 766 distilled / 2,225 read