Two models, no text output

Cloudflare has released Clef and Clef-flash, a pair of decision models that do not generate text at all. In a post on 1 October the company said both are open-weight under Apache 2.0 on Hugging Face and hosted on its Workers AI platform, and that both are API-compatible with Typesafe AI’s Jev System One, the model that created this category a few weeks ago.

A decision model takes an input state and a set of typed questions, and returns a probability for each allowed answer. Cloudflare’s example is a support message: is this urgent, which team should handle it, how severe is the impact. The output is a structured record of probabilities, which calling code can act on without a person reading it.

How it is built

Clef freezes a Qwen3.8-27B backbone and Clef-flash a Qwen3.5-9B one, with rank-256 low-rank adapters and a routing head trained on top. At inference the backbone does a prefill-only pass and the model then scores the valid schema choices in parallel. Nothing is produced token by token, which is where the speed comes from.

A software developer working at a desk with two monitors
The models return typed probabilities for a calling program to act on. Illustrative photo. ThisIsEngineering · pexels · Pexels License

Cloudflare also trained with a Brier loss alongside label-smoothed cross-entropy to calibrate the probabilities, and with a method it calls Reinforcement Learning for Calibrated Decisions, which gives partial credit for adjacent ordinal answers. Two differences from Jev are easy to state: Clef has a vision encoder and can classify images, which Jev cannot today, and it has a 64k context window against Jev’s 32k.

The numbers, and whose they are

Every figure below is Cloudflare’s, measured by Cloudflare. On the ten-row table it published from the Jev Decision Index, a Clef model posts the highest score on seven rows — including 94.20 macro-F1 on BANKING77 against Jev’s 79.74, and 97.73 on a home-appliances set against Jev’s 52.27. Jev leads on When2Call and BRIGHT, and DiffusionGemma Jev leads on PhishNChips.

Fibre optic patch cables plugged into a network switch
Clef and Clef-flash are hosted on Cloudflare's edge network. Illustrative photo. Brett Sayles · pexels · Pexels License

On latency across 43 benchmark runs, Cloudflare reports a median of 209.3ms for Clef and 38.8ms for Clef-flash, against 524.1ms for Jev. A rival called Laya is faster still at 5.8ms median, but Cloudflare’s own table shows it scoring 38.13 where the Clef models score above 98, and its p95 latency is worse than Clef-flash’s. On Typesafe’s own evaluation suite Cloudflare says Clef beat Jev in three of four workflows, losing on agent-trace observability.

The one internal result worth repeating is from Cloudflare’s threat-intelligence team, which uses Clef to categorise domains: 2.2 seconds to fetch, render and classify a site, against 4.7 seconds for gpt-oss-120b, which returned only two classifications.

The fine-tuning business underneath

Alongside the models Cloudflare is launching a reinforcement-learning fine-tuning service, delivered first by its forward-deployed engineering team and later, it says, as self-serve. The pieces are existing products — AI Gateway to collect request data, Workers AI for rollouts, Containers as the RL sandbox — plus a new component called Trainer and the bring-your-own-model work that came out of its Replicate acquisition.

That is the part to watch. Open weights make the model easy to try; the fine-tuning pipeline is what would make customers stay. Jev is two weeks old, Amazon shipped a decision model of its own the same week, and nobody outside the vendors has yet published an independent comparison of any of them.