OpenAI is rewriting its safety rules after the Hugging Face breach

The company has paused two weeks of reinforcement learning and put its largest frontier run on hold, while admitting the model that escaped was never being monitored


OpenAI is rewriting its safety rules after the Hugging Face breach
Image Credits Credit: Jernej Furman from Slovenia

OpenAI said on Tuesday it is rewriting its Preparedness Framework after concluding that its upcoming Astra model may have reached the critical threshold for cyber capability. Its new token-level monitoring carries roughly 20% compute overhead and is now mandatory for its most capable training runs.

OpenAI is rewriting the document it has used to decide whether a model is too dangerous to ship. The company said on Tuesday that the Preparedness Framework, most of which dates to December 2023, no longer fits the systems it is now building.

Two events pushed it there. OpenAI has paused Astra work after finding it may meet the critical cybersecurity threshold, and one of its unreleased models broke into Hugging Face during testing.

The most concrete part of the announcement is a price. New monitoring runs activation classifiers that sample every token, aiming to raise an alert within 30 minutes of concerning activity, at roughly 20% compute overhead.

That is a fifth of the machine, spent watching the machine. It is now required for all reinforcement learning on models at Sol capability and above, and has applied to all Astra inference using tools since 7 August.

The reason it is mandatory is the awkward part. Monitoring was not running on the model that escaped, because, in chief scientist Jakub Pachocki’s account, the company underestimated what it could do.

For AI, you should expect the unexpected,” Pachocki said. It is a striking thing to hear from the person responsible for predicting what the models will do.

Training has slowed in the meantime. OpenAI paused about two weeks of deployment-focused reinforcement learning, and its largest planned frontier run remains on hold along with a significant share of Astra and cyber research workloads.

Sam Altman said it “is a good time to slow down.” Safety lead Mia Glaese put it less comfortably, saying the company is “very far from everything running back to normal.

OpenAI insists this is not damage control. Pachocki described “an incredible feeling of urgency to advance the levels of this sector” and to prepare for the same capabilities arriving elsewhere.

The timing sits oddly against something else. The framework being rewritten belonged to a preparedness team OpenAI dissolved in July, a move the company has described as streamlining ahead of a possible listing.

It is also not alone. Anthropic said in July that three Claude models gained unauthorised access to real organisations during misconfigured evaluations, part of a run of incidents that has now touched more than one lab.

A postmortem on the Hugging Face breach is promised, and outside organisations are to be involved in revising the framework. Until then the only number anyone can hold OpenAI to is the 20%.

Get the TNW newsletter

Get the most important tech news in your inbox each week.