Anthropic’s Dario Amodei speaks with Salesforce’s Marc Benioff among the audience at Dreamforce 2026
Anthropic has named Accenture as the first embedded evaluator of its frontier models, with each company expecting to invest at least $1bn over five years. Anthropic will fund the work directly and says in the same announcement that funding should come from pooled or government sources instead, because neither exists yet. Accenture is already Anthropic’s largest Claude Code deployment, and there are still no standards for what embedded evaluators may access or must report.
Anthropic said on Friday that Accenture will embed evaluators inside the company to red-team its newest models, run alignment assessments and test safeguards. Samantha Oltman and Lynn Doan reported it for Bloomberg, noting the evaluators will hold access comparable to Anthropic’s own staff.
Each company expects to invest at least $1bn over five years. It is the first concrete implementation of a commitment made in Dario Amodei’s essay six days earlier, which TNW covered when it was published.
The funding line is the one to read
Anthropic states plainly that it will fund Accenture’s work directly. It also says that is not the right long-term arrangement, and that funding should come from pooled or government sources.
Its reasoning is that neither of those exists yet. The company argued for them in its own policy framework in June, and in the meantime intends to work with different evaluators under different funding arrangements.
That is a candid thing to publish. It is also a description of a structural problem that an announcement cannot solve, because the evaluator’s invoice still goes to the evaluated.
METR has not been dropped
The essay that prompted this named METR, a nonprofit, as the kind of body it had in mind. Anthropic says it is in dialogue with METR and other nonprofit evaluators to pilot elements of embedded evaluation using their own funding.
So the picture is two-track rather than a substitution. A consultancy paid by Anthropic starts now, and self-funded nonprofits are in discussion.
Which track produces the first critical finding will tell you more than the announcement does.
The commercial relationship is substantial
Accenture and Anthropic already run a joint business group. Around 30,000 Accenture professionals are being trained on Claude, and tens of thousands of its developers use Claude Code in what Anthropic has called its largest ever deployment.
They also fund a Claude centre of excellence inside Accenture and co-develop offerings for regulated industries. Accenture is a customer, a reseller, an implementation partner, and now the evaluator.
Anthropic does not treat this as awkward. It presents Accenture’s enterprise deployment experience as a qualification, on the grounds that understanding how AI is used in practice informs how you assess it.
The case for that is not weak
The work will be led by Faculty, the British AI company Accenture bought in January. Faculty had already worked with leading labs including OpenAI and Anthropic on model safety before the acquisition.
Scale matters too. Embedding a standing team with employee-level access is a staffing problem nonprofits the size of METR cannot solve, and $1bn over five years buys people rather than announcements.
Anthropic has also made the arrangement non-exclusive. More evaluators are promised within weeks, and Accenture is to do the same work for other developers.
What is genuinely unresolved
Anthropic says there are no standards yet for what information embedded evaluators should have access to, or how they should report what they find. That is the admission underneath the money.
An evaluator with employee-level access and no reporting obligation is a governance arrangement defined entirely by the company being examined. Anthropic says independent evaluators make its accountability more verifiable rather than reducing it, and that the safety of its models remains its own responsibility.
The open question is what happens to a finding that would delay a release, inside a firm whose larger business is selling that release to clients.
The same shape keeps appearing
Independent evaluation keeps ending up owned by interested parties. Hugging Face volunteered to audit the AI labs, and Nvidia is buying Hugging Face.
The organisations technically capable of auditing frontier models are the ones the industry wants to acquire or contract. That is the constraint every version of this runs into.
Why it is being built now
External testing has had a bad summer. A single testing vendor sat behind breaches disclosed by OpenAI, Anthropic and Meta, with Google since added.
Anthropic reported three incidents on 30 July in which its models gained unauthorised access to real systems, and said it planned to work with METR on an independent review. It has since resumed the external tests in which its models attacked real companies.
The European detail
Faculty is a London company with a public sector history, bought in January by an Irish-domiciled consultancy, and now the embedded evaluator for an American lab. European AI assurance capability is consolidating into large integrators.
The EU’s own framework leans on independent conformity assessment by bodies without a commercial stake in the product. This is a different model arriving first.
What to watch
Watch whether the nonprofit track materialises and on what terms. Anthropic says more names are coming, and whether any arrives self-funded is the test of whether plurality is real.
Watch who OpenAI picks. Sam Altman said OpenAI would match the commitment, and its choice will show whether a paid consultancy becomes the template or the exception.
Get the TNW newsletter
Get the most important tech news in your inbox each week.