The arrangement
Anthropic said on Thursday that it has picked Faculty, Accenture’s specialist AI unit, as its first embedded evaluator. Faculty will take on “evaluating and red-teaming models, conducting alignment assessments, and testing model safeguards”. Both companies say they expect to invest at least $1bn each over the next five years to build capacity for the work.
The distinguishing feature is access. External evaluators today usually see a model late, through an API, for a fixed window. Embedded evaluators, in Anthropic’s description, work inside the company with access comparable to an employee’s: they can watch models take shape during training, follow the decisions that govern how those models are built and deployed, and talk directly to staff.
Accenture acquired Faculty in January. Its shares rose about 8% in after-hours trading on the announcement.
Why this is the first test of a much-discussed idea
Embedded evaluation has been the concrete proposal inside a month of abstract argument about slowing down. Dario Amodei asked the labs to pace the frontier together; Sam Altman committed OpenAI to embedded safety evaluators on 13 September; OpenAI, Anthropic and Google DeepMind have been discussing a self-regulatory standards body modelled on FINRA.
This is the first of those commitments to name a counterparty and a number.

Anthropic is explicit that the arrangement is not exclusive. The company says it is also in discussions with METR and other non-profits to pilot embedded evaluation funded by those organisations themselves.
The objection
The choice surprised people who expected a safety research organisation. The public debate about embedded evaluators has centred on METR, Redwood Research and Apollo Research — small non-profits whose only business is examining models. Accenture is a consultancy whose business is helping companies and governments deploy AI, including, potentially, Anthropic’s.
Anthropic’s answer is that Accenture brings practical experience of deploying AI at scale in large organisations, and that being a large public company that predates the current AI industry gives it what Anthropic calls functional independence from Anthropic’s own ecosystem.

Critics read the wider scheme differently, arguing that a safety regime the labs design, fund and staff is a way of pre-empting accountability rather than accepting it. Anthropic’s response is that “the safety of our models remains our responsibility” and that evaluators do not reduce that responsibility — they make it more verifiable. Cohere’s chief executive Aidan Gomez, writing this week, called the standards push by the three largest labs a cartel by another name.
What is still missing
Anthropic’s own announcement is candid about the gaps. There are no standards yet for what information an embedded evaluator should have access to, or how it should report what it finds. There is no settled way to fund independent evaluation: Anthropic is paying Accenture directly for now, and says that in the long run the money should come from pooled or government sources rather than from the company being evaluated.
That is the question the arrangement leaves open. An evaluator paid by the lab it evaluates has an obvious problem, and both sides say so. California’s executive order of 18 September and the FRONTIER Act now in the US House both contemplate a statutory version of this role. Whether the voluntary version gets to define the terms first is the thing to watch.