A different way to answer
A group at Stanford’s Hazy Research lab and Nvidia has released CLM-8B, an open-weights model that does not generate text at all. Given a situation and a list of things it could do, it produces an embedding for each and returns the action whose embedding sits closest to the state.
That is the whole idea. Where a language model writes out a choice token by token, a contrastive language model treats the choice as a matching problem. The authors are Jacky Kwok, Hangoo Kang, Tarun Suresh, Jon Saad-Falcon, Marco Pavone, Christopher Ré and Azalia Mirhoseini. Kwok announced it on 23 September.
It is the second model of this shape to appear this month, after TypeSafe AI’s Jev, which opened early access on 15 September. CLM-8B is the open one.
What is actually trained
The 8B in the name is a frozen Qwen3-8B encoder. What the team trained is two small projection heads — one for states, one for actions — sitting on top of it, using a bidirectional InfoNCE contrastive objective that pulls a correct state-action pair together and pushes competing actions apart.

Training ran in three stages: about 60 million Nemotron question-and-answer pairs, then about 30 million synthetic hard negatives, then about a million agent trajectories.
Because states and actions are encoded separately, an action’s embedding can be computed once and reused. That is where the speed comes from: the candidate set does not have to be re-read on every decision.
The numbers, and who produced them
All of the figures below are the authors’ own, published with the model.
Zero-shot, they report performance “on par with Jev on computer-use, gaming and tool-calling tasks, with up to 9× lower latency”. VentureBeat’s write-up gives the underlying comparisons: 95.2 per cent on the BFCL v4 tool-calling benchmark against Jev’s 99.2, and 26 of 30 WikiRacing tasks completed against Jev’s 30. The trade is accuracy for latency, and the model card says so.

Fine-tuned as a verifier — scoring candidate solutions rather than choosing tools — they report 81.6 per cent on DeepSWE and 87.6 per cent on Terminal-Bench 2.1, which they describe as state of the art for that role, at four to six times lower latency. None of these results has been reproduced independently yet.
Why it matters
Routing, triage, ranking and output selection are calls software makes constantly, and today most of them are made by asking a large language model to write an answer and then parsing it. A model that returns a choice directly removes both the generation cost and the parsing step.
Jev showed there is a market for that. CLM-8B puts the same shape under an Apache 2.0 licence, with the code on GitHub and a playground alongside it, which means the comparison can now be run by anyone rather than reported by either vendor.
The team says a multimodal CLM-35B, trained on a wider set, is due in early October.