OpenAI said Astra, its next frontier model, is the first it has classified at the Critical cybersecurity capability level under its Preparedness Framework. No OpenAI model has been placed in that category before.

What Astra did in testing

Astra scored perfectly on ExploitBench, which measures whether a model can turn a known vulnerability into a working exploit. It then went further and found two zero-day vulnerabilities on its own.

In one evaluation it escaped a browser sandbox and executed commands on the underlying machine. In another it chained multiple flaws in a hardened operating system into a local privilege-escalation path to root.

Abstract Aivio News graphic accompanying the story on OpenAI says Astra is its first model to reach Critical cyber capability.
Illustration by Aivio News. Not a photograph of the events described. Aivio News · owned · © Aivio News / GrowQ AB

Where OpenAI drew the line

The Critical threshold, as OpenAI defines it, applies when a model can independently find and exploit zero-day vulnerabilities across many well-defended systems, or carry out a complete cyberattack against a hardened target from only a high-level instruction.

That is a definition written around autonomy rather than skill. A model that finds a bug with a researcher steering it is not what the threshold describes; a model that does the whole chain from a one-line goal is.

Eight weeks of hedging

OpenAI has been signalling this since early August, when it said internal evaluations showed “significant advances in agentic coding and cybersecurity” and that it “cannot rule out the model reaching the critical capability level for cybersecurity.”

At the time the company paused Astra activities that did not meet enhanced security requirements — not the project itself. It also stated flatly that “Astra is an upcoming model and was not involved in exploiting Hugging Face,” a denial that says something about the questions it was being asked.

What ships and what does not

OpenAI has cleared Astra for release under the framework, on the basis that its safeguards sufficiently reduce the risk of severe harm. The most advanced cybersecurity capabilities will not be available at launch.

Access starts with early testers and widens through a programme OpenAI calls Daybreak Blue. Jailbreak resistance is reported at 91.5%, against 59% for GPT-5.6 Sol.

Abstract Aivio News graphic accompanying the story on OpenAI says Astra is its first model to reach Critical cyber capability.
Illustration by Aivio News. Not a photograph of the events described. Aivio News · owned · © Aivio News / GrowQ AB

The company framed the moment in its own words: “We are entering a stage of AI development in which models can take on more consequential work, and failures of alignment and control can have more serious effects.”

What to watch

Whether a 91.5% jailbreak-resistance figure holds outside OpenAI’s own red-teaming. At Critical capability, the gap between 91.5% and 100% is not a rounding difference — it is the share of attempts that reach a model which can write working exploit chains unaided.