One API call for audio, video, images and text
Cloudflare has added a multimodal member to its Clef family of decision models and cut the price of the cheapest one by more than half. Clef-omni takes audio as wav or mp3 and video as mp4 or webm, alongside images and text, and processes all of it in a single pipeline and a single API call. The company published the details on Thursday in a post by Michelle Chen.
The model is built on Alibaba’s Qwen3-Omni-30B-A3B-Instruct, a mixture-of-experts model. Cloudflare says it kept the comprehension backbone and discarded the text-to-speech output components, then trained with the Qwen3 backbone frozen, low-rank adapters, and a label-smoothed cross-entropy loss plus Brier score calibration. Like the rest of the family it generates no output tokens at all; it returns a probability over a fixed set of options.
What it costs and how fast it is
Clef-omni is priced at $0.15 per million input tokens. Cloudflare gives median latencies of roughly 130 milliseconds for text alone, about 150 milliseconds with images, a few hundred milliseconds for audio, and about 1.5 seconds for a 21-second video with sound. The weights are published on Hugging Face as Cloudflare/clef-omni, and the model id on Workers AI is @cf/cloudflare/clef-omni.

The sharper change is to Clef-flash, which drops from $0.09 to $0.038 per million input tokens — cheaper, Cloudflare notes, than Jev, the TypeSafe AI model that started the category. The price comes with a trade: the hosted context window falls from 64,000 tokens to 24,000. Cloudflare says only 0.24% of requests exceed 24,000 input tokens, and that the Hugging Face weights, trained for 256,000, are unchanged.
A serving change, not a new model
The base Clef model keeps its $0.24 per million input tokens and its 64,000-token window, but got faster: Cloudflare moved it to SGLang, contributing support in pull request #42721 and shipping on SGLang 0.5.22. Its published figures show median latency falling from 262ms to 152ms on roughly 800 tokens, from 616ms to 305ms at about 3,400 tokens, and from 2,721ms to 1,635ms at about 16,000. The weights did not change.

“We heard your feedback — you want a decision model affordable enough to incorporate into any workflow,” the company wrote.
The benchmark table is Cloudflare’s
The post publishes scores on ten public benchmarks and five internal workflow evaluations, all run by Cloudflare. On its figures Clef-omni reaches 94.8 macro-F1 on BANKING77 against 79.74 for Jev, and 97.7 on CLINC150+OOS against 89.27. Jev leads on others, including When2Call, where Cloudflare puts Jev at 80.97 and Clef-omni at 63.3. No independent evaluation of either has been published.
Cloudflare says Clef is API-compatible with Jev, so switching is a model-id change, and that its own teams use the family to close spam issues in a public docs repository, moderate plugin libraries for phishing, scan for personal identifiers and detect malicious domains. The Workers AI changelog carries the same dates. What is worth watching is whether the 24,000-token window turns out to bite the 0.24% harder than the average suggests.