Amazon’s Strands Labs has released Strands Decider 2B, an open-source model that does not generate text at all. Given a set of predefined options, it scores them and returns the best one with a confidence value. That is the entire output.
It is fast because of what it does not do. Strands Labs reports a median latency of about 115 milliseconds on a local Nvidia RTX 3090 and about 153 milliseconds on an M3 MacBook.
How it is built
The model takes a pre-trained Qwen3.5-2B torso, removes the language-modelling head and replaces it with a pointer head that scores the hidden states at each option position against the position of an answer token. The pointer head is just over a million parameters. Fine-tuning uses a rank-16 LoRA adapter.

Strands Labs says it ranks third of 33 models in its two-billion-parameter class on JevBench for accuracy and calibration combined — calibration being measured by Brier score, which asks whether the confidence value means anything rather than whether the answer is right. The weights are on Hugging Face and the training data, training scripts and the full iteration history are on GitHub.
The model class is a month old
Decision models did not exist as a category in September. TypeSafe created one with Jev, named after the economist William Stanley Jevons and his argument that a falling cost can raise demand rather than lower it. Dozens of imitators have followed.

Marc Brooker, the Amazon engineer who led the work, framed the appeal in operational terms: a workflow step that can be structured to be more reliable, with lower latency and potentially lower cost. TypeSafe’s chief executive was less generous about the field it started, telling TechCrunch the current batch seems more like machine-learning people wanting to implement an interesting architecture than a team dedicated to making intelligence useful.
Why this matters past the novelty
Most agent frameworks route by asking a large model to emit a word and then parsing it. That is expensive, slow and silently unreliable, because a model that writes “option B” gives no usable measure of how sure it was. A dedicated head that returns a calibrated score over a fixed option set removes both problems from the step where agents most often go wrong.
What to watch
Whether JevBench holds up as the measure. A benchmark created alongside a product by the company that made the category is the obvious thing for a rival to dispute, and so far no one has.