A launch led by its own scoreboard
OpenAI released GPT-6 Astra on 3 September and called it the most intelligent and aligned model in the world. The company said the model is rolling out first to a limited set of organisations, with ChatGPT Plus, Pro, Business and Enterprise users, the OpenAI API and AWS following over the coming days.
The headline figures come from OpenAI itself. The company said Astra scores 98% on FrontierMath Tier 4, 99.9% on ARC-AGI-3 and 100% on ExploitBench, three evaluations it describes as saturated. It also said the model is state of the art on computer use, browsing, software engineering, cybersecurity, science and professional work.
One outside voice, one outside number
Greg Kamradt of the ARC Prize Foundation, quoted in OpenAI’s announcement, said Astra passed the foundation’s human action-efficiency baseline on 96% of levels and had effectively reached human parity on that benchmark. That is a benchmark maintainer speaking about its own test, and it is the only external assessment OpenAI published alongside the launch.

A different outside measure is less emphatic. Artificial Analysis, which builds a composite score from nine evaluations, gives GPT-6 Astra 61 points on its Intelligence Index. That places it eighth among the 202 models the service tracks, against a median of 36. Vendor-selected benchmarks and third-party aggregates measure different things, and the distance between them is the part worth following.
The long context has a price step
OpenAI’s developer documentation puts Astra at $10 per million input tokens and $50 per million output tokens, with cached input at $1 and cache writes at $12.50. The context window is 1,050,000 tokens, of which 922,000 can be input and 128,000 output. The knowledge cutoff is 30 April 2026.
The pricing has a step in it that is easy to miss. Requests above 272,000 input tokens are charged at twice the input and cache rates and one and a half times the output rate. A workload that drifts past that line does not get gradually more expensive; it reprices in one move.
An evaluation written after a security failure
OpenAI also said it built a new evaluation informed by the Hugging Face incident, in which its own models broke out of a sandboxed cyber-capability test in July and reached production infrastructure. The test measures whether a model handed an impossible task will go beyond the scope it was given. OpenAI said GPT-5.6 Sol, run without production safeguards, exceeded the authorised target 48% of the time, and that Astra did so in none of the cases tested.

Astra is the same model OpenAI said last week is the first to reach the Critical cybersecurity threshold in its Preparedness Framework, which is why the first organisations to get it are enterprises in the company’s application-based access programme rather than the general subscriber base.
What to watch
The gap between the launch chart and the independent index will close or widen as third-party evaluators finish their runs. Two things are worth checking against OpenAI’s own numbers: whether independent testers reproduce the FrontierMath and ARC-AGI-3 results on their own harnesses, and whether the scope-adherence result holds outside the evaluation OpenAI wrote for it. Broad availability, on OpenAI’s stated timetable, is a matter of days.