A model that scores options instead of writing sentences
Microsoft has released Microsoft-Decision-1, a model that takes a question with a fixed set of answers and returns a calibrated probability for each one. It does not generate prose. Announced on Thursday by Achint Srivastava, a vice-president of software engineering in the Office of the CTO, it is Microsoft’s entry in a category that did not exist in June.
The model is a post-trained version of Alibaba’s open-weight Qwen3.5-9B, tuned for single-pass decision scoring. Microsoft says it will later rebase the same system on other models, including its own MAI family and OpenAI’s. It handles yes/no questions, multiple choice and rating scales, and rubric-based grading of other models’ responses and of agent actions.
The numbers are Microsoft’s own
Microsoft says Decision-1 “achieved the highest accuracy in our 36-benchmark comparison”, a set covering nearly 150,000 questions it says were kept blind from training, across routing, ranking, long context, multilingual and out-of-distribution tasks, reasoning and safety. The comparison included several of the top models from the JevBench leaderboard. These are vendor-run results, and no independent evaluation has been published.

On stability, the company reports that across eight kinds of input perturbation the model’s decision flipped 1.3% of the time on average, and that paraphrasing the option descriptions or reordering the options caused no flips at all. On safety it reports testing across 5,250 requests and 11 benchmarks covering harmful content, jailbreaks and prompt injection, without giving a score.
Price is the argument
Decision-1 costs $0.042 per million input tokens. Output tokens are free, which follows from the design: there are no output tokens to speak of. Microsoft puts the model’s median latency at roughly 35 times faster than GPT-6 Sol, and 2.5 times faster than the runner-up in its comparison, H2O-Lightning-4B v1.1.

It is available through Microsoft Foundry and on OpenRouter, with documentation on Microsoft Learn. Microsoft has not said what licence the weights carry, or whether they are published at all.
A crowded four months
The category started with Jev, from TypeSafe AI, in mid-September, and has filled up since: OpenAI shipped a Decisions API this month, Cloudflare has an open-weight family called Clef, and Amazon released a clone. The Decoder noted that Cloudflare’s Clef models, which are also built on Qwen, were left out of Microsoft’s comparison.
What to watch is the rebasing. Microsoft has said in its own announcement that Qwen3.5-9B is a starting point rather than the destination, which makes Decision-1 the rare frontier-lab product that names an open Chinese model as its current foundation.