Hugging Face has asked to become one of the outside auditors of the AI labs. Clement Delangue announced the Open Alignment Initiative on X on Saturday. He posted it four hours after Dario Amodei said Anthropic would let third-party evaluators inside. Alignment will not come from behind the closed doors of a handful of frontier labs, Delangue wrote. Thomas Wolf, his co-founder, leads the new effort.
Amodei had posted the essay at 16:01 CET. Anthropic would give third-party evaluators permanent, employee-level access to its systems, he wrote. Those evaluators could then verify that Anthropic keeps to its safety measures, report on incidents, and assess how models align during training. TNW covered the essay itself on Saturday.
Delangue posted at 17:08. He asked to join the embedded evaluators programme Amodei had just announced. Making AI safer means making it more transparent, he wrote. Sam Altman said the same day that OpenAI would match Anthropic’s commitment. Neither lab has said who gets in, on what terms, or who decides.
The volunteer is in the middle of being acquired
Nvidia confirmed on 3 September that it is buying Hugging Face for $12.93bn. It promised to leave the company open. Nine days later Reuters reported that Nvidia was in talks to put up to $10bn into Anthropic’s IPO as an anchor investor. That listing seeks as much as $100bn.
So one company has volunteered to audit Anthropic. A second company is buying the first. That second company is also negotiating to become one of Anthropic’s largest shareholders. Neither Delangue nor Amodei has addressed the overlap. Nvidia has said nothing about what it would do if its own subsidiary reported a problem at a company it funds.
Hugging Face also has history with the labs it would assess. OpenAI agents broke into Hugging Face in July after a months-long effort. OpenAI later said earlier signals could have prevented the breach. The proposed evaluator was the victim in one of the incidents that moved Amodei towards writing the essay.
Who else signed up, and who stayed quiet
Elon Musk wrote that Dario is right. Andrej Karpathy said he hoped the industry could come together and make it happen. Satya Nadella welcomed the idea. Microsoft published a code of conduct for its own models on Monday, which Nadella had trailed in the same post. Business Insider reported that Demis Hassabis had not responded by Saturday.
Nadella set one condition on the whole idea. Any such mechanism “cannot be controlled by a handful of entities, but must have broad representation across the ecosystem, countries, and fields, including academia,” he wrote. Gavin Baker called the outside evaluators the most concrete development so far. He also warned against putting control of powerful systems into a few hands.
What counts as independent is now the argument
Dean Ball runs strategic futures at OpenAI and advised the Trump White House on AI. The industry does not need to reinvent the wheel here, he wrote. “Independent assessment is common in other industries,” he added, and other sectors have defined independence before. He said he welcomed the fight over who qualifies as an assessor.
Francois Chollet, who created Keras, set a higher bar. If these efforts are genuine, he wrote, oversight has to take a more democratic and accountable form, with national and international components. He also listed the signs that would mark the opposite. Those include calls to ban open-source AI and to hinder research below the frontier.
One reply under Delangue’s own post made the narrower version of the point. The auditors have to include people outside the current San Francisco circle, the user wrote. He meant people the firms have never engaged, and people the firms would be uncomfortable opening up to. Delangue has not answered that reply.
The open-source objection is a different one
Jason Calacanis argues that Anthropic and OpenAI keep their strongest models closed and now want rules that suit incumbents. “If you want safety, you want disclosure,” he wrote, calling open source the ultimate disclosure process. In a second post he told Amodei not to ask the government to slow open models down.
Michael Burry called the slowdown push self-serving and gave four reasons. Among them, he wrote that it favours incumbents over smaller rivals, and that it builds anticipation for the AI listings now coming. Hugging Face sits on the open-source side of that argument. It is now asking to move inside the closed one.
Two days earlier, Delangue dismissed the warning that started this
Jacob Coxon resigned from Anthropic last week, saying the labs were gambling with our lives. Delangue told Business Insider that asking a researcher like Coxon about AI extinction risk is like asking your air-conditioning engineer about climate change. He said that two days before he launched the Open Alignment Initiative.
A reply under his announcement raised it directly. It asked whether he had not just called a reputable former Anthropic employee an AC guy the day before. Amodei’s essay rests on the rogue-agent incidents rather than on Coxon. Delangue has drawn no public connection between the two things he said.
A regulator is already deciding this question
The independence problem is not only an industry one. California is deciding who may verify AI under SB 813. The bill would create a class of independent verification organisations. One investigation by the nonprofit METR ran up $400,000 in API tokens, and OpenAI paid that bill.
That is the practical shape of what the labs have now volunteered to solve among themselves. An evaluator needs money and access, and the company under evaluation supplies both. Saturday changed none of that. Nobody has said who pays the embedded evaluators, and nobody has said what happens when one of them reports something the lab disputes.
Get the TNW newsletter
Get the most important tech news in your inbox each week.