What it returns
OpenAI has opened a Decisions API in public beta. It takes text, images or both, and returns a typed answer rather than prose.
Three answer types are documented. A predicate returns a probability between 0 and 1 that some condition holds. A choice returns one of the values you supplied, with per-option probabilities and a confidence. A score returns a probability-weighted average over level indices, again with probabilities and a confidence. The code examples also handle a refusal.
That is the whole surface. There is no free text, no tool call, no streaming paragraph.
The price is the design
Input costs $0.10 per million tokens, and OpenAI says you pay for input tokens only: no cache-read charge, no cache-write charge, no output-token charge. Regional processing premiums and long-context multipliers still apply.

Free output is less generous than it sounds, and more interesting. A probability and a confidence score are a handful of tokens; charging for them would be rounding error. What the pricing actually signals is that OpenAI expects this endpoint to be called at volumes where per-request overhead is the thing that matters.
The documentation claims the endpoint answers about ten times faster than the Responses API, without giving absolute latencies. Only gpt-6-luna is available on it so far.
The constraints worth reading
Images have to arrive as inline base64 data URLs. Hosted HTTP or HTTPS image URLs and file_id inputs are not supported on this endpoint, which is a real constraint for anyone whose images already live in object storage.
On the compliance side, the endpoint supports zero data retention and HIPAA use for eligible customers, with data residency and regional processing available in the United States and in Europe. OpenAI says it expects general availability in the coming weeks.

The tier change alongside it
OpenAI also cut its paid API tiers from five to three, according to its rate-limit documentation: Build, Launch and Grow, with monthly usage limits of $500, $5,000 and $200,000. Organisations move up automatically once their total credit purchases reach the next threshold.
Why a decision endpoint at all
Most production uses of a language model are not writing. They are classification, routing, extraction, moderation and scoring — jobs where the useful output is one field and the text around it is waste the caller has to parse and the provider has to bill for.
The Decoder reads the launch as a response to the decision-model trend that smaller labs started earlier this year. That is interpretation rather than an OpenAI statement, but the shape fits: Anthropic cut its small-model price the same day, and both moves point at the same tier of work.
What to watch
Whether more models arrive on the endpoint, whether calibration figures are published for the probabilities it returns, and whether the general availability date holds. A probability that is not calibrated is a number with a decimal point, and a routing layer built on one will be confidently wrong at scale.