Earlier than the week before launch
OpenAI published a document on Tuesday setting out how it wants outside organisations to assess its models. The company describes it as priorities and principles for rigorous, secure, and independent third-party AI safety assessments of frontier models and safeguards.
The substance is a change of timing. Until now, external groups have been brought in shortly before a model ships, to run evaluations on something close to the finished article. OpenAI now says assessors can work during training, evaluation and deployment, Bloomberg reported on Tuesday.
That matters because a pre-launch evaluation arrives after every expensive decision has been made. An assessor who can see the training run can ask about choices that are still reversible.
Who might do it
OpenAI says it is in talks with groups including METR and Redwood Research, according to The Next Web, which reported that no formal partnership has been confirmed and no timeline disclosed. Both organisations previously looked into an incident in which an OpenAI model reached systems it was not meant to reach on Hugging Face.
Lama Ahmad, who leads OpenAI’s engagement with external safety experts, set out the conditions the company thinks such reviews need: independence mechanisms, scientific rigour, security practices and clear responsibilities. OpenAI also says some assessors may work from its offices for the most sensitive stages.

Less specific than what Altman promised
The Next Web noted that the announcement is looser than the commitment Sam Altman made ten days earlier, when he described independent evaluators with desks, badges and laptops, and with the right to publish. Tuesday’s document mentions office access for the most sensitive work but does not set out access conditions, publication rights or a named partner.
The gap between those two statements is the whole question. An evaluator with a badge and no right to publish is a consultant. An evaluator with a right to publish is a check.
The comparison next door
OpenAI is not the only lab rearranging this. Last week Anthropic named Accenture as an outside evaluator, in an arrangement in which each side is putting up $1bn, while acknowledging that paying your own examiner creates an incentive problem.

Europe has a different model for the same job. Under conformity assessment, the body that certifies a product is accredited by a regulator rather than retained by the manufacturer, and it does not depend on the manufacturer for its next contract. Nothing announced by either American lab so far works that way.
What to watch
Three things will show whether this is a governance change or a communications one: whether a partner is actually named, whether that partner can publish findings without OpenAI’s approval, and whether assessors get access to training decisions rather than a seat in the room where finished models are demonstrated.