What Clef actually does
Cloudflare has spent years insisting it is infrastructure, not a lab. On October 1, during its Birthday Week, it crossed that line: Clef and Clef-flash are the first models trained by the company's own Workers AI team, and both are already live on Workers AI and open-sourced on Hugging Face under Apache 2.0.
But these are not chatbots and not image generators. They belong to a category that got its name three weeks ago, when Typesafe AI launched Jev as a "System One" decision model. A decision model reads an input state — a support ticket, a web page, an agent's trace — plus a set of typed questions, and returns a probability for every allowed answer. Route the ticket. Block the request. Escalate. There is no free-form output to parse and no reasoning tokens to wait for.
Cloudflare's numbers, read carefully
The spec sheet: Clef packs 27 billion parameters on a Qwen 3.8-27B base; Clef-flash is the 9B sibling on Qwen 3.5-9B. Both carry a 64k-token context window and a vision encoder that takes up to four images per request. Pricing is $0.24 per million input tokens for Clef and $0.09 for Clef-flash, with output tokens unbilled entirely, which makes sense when there is no output to bill.
Then the benchmarks, all from Cloudflare itself, so read accordingly. Across 43 of its own runs, Cloudflare says Clef answers at a 209ms median against Jev's 524ms, and Clef-flash at 39ms. On BANKING77 classification, Clef posts a 94.20 macro-F1 versus Jev's 79.74. Cloudflare also claims API compatibility with Jev, meaning a team built on Jev could theoretically swap endpoints without rewriting integrations. Theoretically is doing work in that sentence: API-compatible does not mean behavior-identical, and nobody outside Cloudflare has published independent numbers yet.
The models are the wedge
The release that matters may not be the models. Alongside them, Cloudflare debuted a reinforcement learning fine-tuning service for Clef: capture your decision data, tune the model, redeploy, all on Cloudflare's stack. It launches through design partners first, with a self-serve version to follow.
That is the tell. The decision model is the wedge; the platform is the product. Agents need thousands of cheap, fast, calibrated judgments a day — route, approve, escalate — and paying full LLM prices and latency for each one is a tax nobody wants. In three weeks the category has gone from one product (Jev, September 15) to four contenders: OpenAI mentioned its Luna model and a Decisions API at DevDay on September 30, and Amazon shipped the open-source Strands Decider 2B the same day as Clef. OpenClaw creator Peter Steinberger's reaction — "Never seen an idea spreading so fast" — is unusually honest for a launch-day quote, because it is observably true.
For agent builders, the practical read: the expensive-LLM-for-everything architecture now has a cheaper routing layer, and it is open source. For everyone else, the read is simpler. Cloudflare trained its first models not to talk, but to decide. That choice says more about where the industry thinks the money is than any keynote would.
Sources
- [1] Cloudflare Blog — “Introducing Clef: Cloudflare’s decision models”Read source
- [2] Cloudflare Developers — Workers AI changelogRead source
- [3] MarkTechPost — Cloudflare releases Clef and Clef-flashRead source
- [4] Dealroom — Cloudflare launches Clef as decision models spreadRead source
- [5] Crypto Briefing — Cloudflare Clef decision models on Workers AIRead source
- [6] Startup Fortune — Amazon releases a Jev rivalRead source