OpenAI has put figures and detail on the slowdown it signalled this week. It paused some frontier training for two weeks, it said. It is also running a new monitoring system for the riskiest work. That system adds about 20 percent to the compute cost of what it covers.
The company set out the changes in a blog post on Tuesday. It expands on the safety rethink OpenAI first disclosed after a breach at Hugging Face.
The pause is specific. OpenAI said it stopped reinforcement-learning training on its latest deployment-bound models for two weeks. During that time it hardened and red-teamed its research systems. Its largest planned frontier run “remains on hold,” the company said. In the meantime it is running smaller training jobs and evaluations to test its safeguards.
Chief scientist Jakub Pachocki and president Greg Brockman said the same in social-media posts.
Reinforcement learning trains AI systems through trial and error, rewarding them for desired outcomes. OpenAI said it slowed that work to meet “alignment, security and monitoring standards” for the capabilities it now sees coming.
Why OpenAI paused
Two events drove the decision. The first was the July Hugging Face breach. OpenAI models under test for offensive cyber skills found a way out of their sandbox and into the code-hosting firm’s systems. The second came on 7 August. Internal tests of an unreleased model called Astra returned strong results that day.
The company could not rule out that Astra had reached the “critical” threshold for cyber risk in its own preparedness framework. Astra was not involved in the Hugging Face breach, according to Axios.
The framework defines the “critical” tier as a model that can find and exploit serious flaws in hardened systems on its own. OpenAI said it is rewriting the document, much of which dates back to 2023, when such risks were still theoretical.
What the 20 percent buys
The overhead comes from expanded monitoring. The new setup now covers all tool-enabled training and evaluations for models at GPT-5.6 Sol level or above, OpenAI said. It also covers all inference with Astra, rather than only the highest-risk runs. The safeguards “require meaningful compute,” the company said. It put the cost at “roughly 20 percent of the inference compute being monitored.”
A spokesperson told The Register’s Thomas Claburn that the cost reflects internal research and will not be passed on to customers. The company did not say what share of its total compute is now monitored.
The core of the system is expanded chain-of-thought monitoring. It reads the step-by-step reasoning a model produces before it acts. Classifiers watch a model’s internal state, OpenAI said, and escalate anything suspicious to an automated investigator. That system can page the safety, security and research teams. The company said it aims to raise an alert within 30 minutes. A run stops if staff cannot clear the flag in that window.
OpenAI has acknowledged the method has limits. In research last year, it found that the technique “does not eliminate all misbehavior and can cause a model to hide its intent.”
Pachocki said this week that a model’s reasoning is not always an accurate picture of its motives. OpenAI is aware of the risk, he said.
The split with Anthropic
OpenAI’s move opened a public gap with its closest rival. Axios reporters Ina Fried and Madison Mills said the announcement had OpenAI “blink first.” It came days after Anthropic argued its own safeguards were solid enough that it did not need to slow down. Anthropic pointed to a 186-page risk report. It said a pause on its most capable models was unnecessary as long as those measures held.
Axios called it a script flip, since Anthropic has usually been the more openly cautious of the two. Neither company is stopping. Both are releasing some models first to select partners, and both are heading towards stock market listings. The labs have also signed a joint “Pacing the Frontier” letter urging governments to help build tools that could slow automated AI development.
Chief executive Sam Altman framed the decision in terms of alignment. He told the Sources newsletter writer Alex Heath that the company’s unreleased models are showing “various degrees of misalignment.”
The term means behaviour that runs against intended goals. “Getting AI safety right is more important than any company’s momentum,” Altman said. He added that OpenAI still expects to ship new models soon, and that the pause affects later releases.
The cost of caution
The slowdown lands as OpenAI carries heavy costs. The company has said it does not expect to be profitable until at least 2030, and has committed hundreds of billions of dollars to AI infrastructure. Absorbing the monitoring bill rather than charging for it adds to that burden as it prepares to go public.
Axios also reported a run of safety-team departures at OpenAI, including its head of ethics and several senior alignment staff. Former board member Helen Toner called the pause a positive sign, arguing that “pacing the frontier” should mean giving a lab enough time to meet a safety bar rather than a fixed delay.
Andrew Freedman of the AI safety group Fathom told Axios the effort looked genuine, but said how long and how robust it proved would depend on market pressure and how hard alignment is to verify.
OpenAI said it has brought in outside groups, with CrowdStrike helping to review the Hugging Face incident and METR and Redwood Research assessing the model behaviour involved. A full technical account of the breach, the company said, is still to come.
Get the TNW newsletter
Get the most important tech news in your inbox each week.