A model that cannot write
Perplexity released pplx-decider-v1-27b on Thursday, an open-weight model published on Hugging Face under the Apache 2.0 licence and fine-tuned from Qwen3.8-27B. It is multimodal, taking text and images, and it does not generate prose. It returns a selected choice from a fixed set together with calibrated probabilities.
That is the defining property of the category. A decision model answers in one of a few typed shapes — yes or no, one of n options, an ordered score — and because the output space is closed, the answer can be checked rather than parsed. For an agent deciding whether a retrieved passage supports a claim, or whether a support ticket is a refund request, that is the entire requirement.
What the benchmarks say, and who ran them
On Perplexity’s own panel of eleven benchmarks and 7,210 samples, pplx-decider scored 85.71 per cent overall against 84.51 per cent for Jev, the TypeSafe AI model that opened the category in September. The model card reports 90.60 per cent on TabFact, 89.20 on Circa, 88.80 on RAGTruth, 84.18 on FinancialPhraseBank and 83.30 on WinoGrande, and a 74.76 per cent overall for the base Qwen3.8-27B it was fine-tuned from.

The margin over Jev is 1.2 points on a panel Perplexity selected, and the breakdown is not uniform: by AI Weekly’s count Jev scores higher on six of the eleven tests, with Perplexity ahead on FinancialPhraseBank, RAGTruth, TabFact and Circa. The widest gap is RAGTruth, where Perplexity reports 88.80 against 77.27.
A one-point lead on a vendor-chosen panel is not a capability claim worth much on its own. The weights being on Hugging Face under Apache 2.0 is the part that can be checked by anyone with a GPU.
Three in two days
Amazon opened Strands Decider, a two-billion-parameter model, on Wednesday. Cloudflare opened two Clef models the same week. Perplexity’s is the largest of the three at 27 billion parameters, and the only one fine-tuned from an existing open base rather than trained for the purpose.

The speed of that convergence is the actual story. Agent systems make a very large number of small classification calls, each of which currently runs through a general-purpose model that was built to write and is being asked to pick. Charging for generation tokens that are never generated is an obvious inefficiency, and the market has found it all at once.
What to watch is whether an independent evaluator publishes a panel none of these vendors chose. Four decision models now exist and every comparison between them so far has been run by one of their authors.