Paperwork before the training run
OpenAI published a set of guidelines on Sunday for what it thinks should be documented before a frontier reinforcement learning training run is allowed to continue. The framing is blunt: “we believe we are entering a new era in which structured safety documentation should be required before continuing any frontier reinforcement learning training run.”
The model it borrows is the safety case — the structured, evidence-based argument about risk used in aviation and nuclear power. OpenAI calls this “an aspirational north star we are building towards,” and concedes the analogy is imperfect because of “the emergent complexity at each new level of AI capability.”
Three technical layers, and one rule about chains of thought
The technical half of the document covers alignment training, containment and monitoring: train the model not to act against intent, make it hard to break out if it does, and catch it if it tries.
Under alignment, OpenAI proposes using agents to find and fix broken reinforcement learning environments where a reward hack could score highly, tuning graders to penalise exploitation, and backtesting alignment evaluations against previous incidents so they are not merely overfitted to them. It also proposes tracking whether a model recognises it is being tested, with “specific blocking thresholds.”

One line is unusually specific: do not let automated graders see the chain-of-thought during reinforcement learning, because doing so risks models “evolving to evade chain-of-thought monitors.” That is a deliberate decision to leave capability on the table to keep the model’s reasoning readable.
Containment covers hardened sandboxes, red-teaming the sandbox using checkpoints from actual frontier runs, closing high-bandwidth side channels between samples, and keeping agent transcripts in write-once storage so an investigation cannot be edited after the fact. Monitoring asks for high recall on past incidents, fresh evaluation data so recall figures do not go stale, and a response rule: page the on-call during the day, and auto-pause runs on unacknowledged alerts at night.
Who gets to stop a run
The operational half is where the governance is. OpenAI says that after a safety case is drafted, someone from another team should write a dissent — a pre-mortem — to find holes in it. It says the case should be approved by senior leaders who “should each have the ability to veto the run,” naming a research org lead or VP, the Head of Safety and the Chief Scientist.

It goes further on accountability: the senior leader responsible for a training run should answer for the safety case and any incident response “including as part of performance reviews.” Escalation should run through defined misalignment severity levels, with an on-call able to page executives — the document names the CEO — once a run hits a given severity. Safety features should fail closed: it should not be possible to start a run without monitoring enabled, or to disable the monitor from inside the training run itself.
Auditors, the document says, should get enough access to verify the claims are sound.
What it is, and what it is not
OpenAI says these recommendations are “in the process of being implemented” internally and that its practices will keep evolving “over the coming weeks.” Nothing here is a commitment to an outside party, and no regulator is named.
The context is unavoidable: the document went up the day before OpenAI confirmed it would not release GPT-6.1 Astra, and the day before it named four Australian government bodies its models reached during internal training. What to watch is whether any of this becomes auditable by someone who does not work at OpenAI.