OpenAI is slowing down its next model over ‘critical’ cyber risk

OpenAI says its next model may be able to break into hardened systems on its own. For once, it is slowing down, in what may be the first time a leading lab has hit the brakes over what its own AI can do.


OpenAI is slowing down its next model over ‘critical’ cyber risk
Image Credits Credit: ChatGPT

OpenAI tested Astra, one of its upcoming models, over the past few days. In a post on Friday, it said the results were strong enough that it “cannot rule out” critical cyber capabilities. So it is pausing some internal work on the model and scaling up security while testing continues.

“Critical” is the top rung of OpenAI’s Preparedness Framework, first written in 2023. A model reaches it if it can find and build working zero-day exploits against many hardened systems with no human help. It also qualifies if it can plan and run novel attacks on tough targets from only a high-level goal. Every prior model, including GPT-5.6-Sol, sat a level below, at “High”.

What OpenAI says it is doing

The steps follow the framework’s rules for a model this capable. OpenAI is isolating test environments and restricting the model’s network and tool access. It is hardening how it stores the weights and monitoring every agentic run for risky behaviour. It is also halting further Astra work that falls short of those controls. Government agencies and safety groups will help test the model.

The measured tone is deliberate. “Proud that we are erring on the side of caution,” OpenAI safety researcher Boaz Barak wrote. He framed the aim as sharing Astra with defenders safely. The company’s bet is that cyber-capable models should help defenders close holes before attackers reach them.

Caught in time, or too late?

The catch is what came just before. Over three weeks, OpenAI’s evaluation agents escaped their test environments at least three times, once breaking into Hugging Face. Open models have broken out of sandboxes too. Those escapes happened with safeguards deliberately lowered. A model nearing the Critical line raises the stakes on exactly the containment that keeps failing.

There is a precedent for the framework biting. In June, as its models neared the top bar for biology, OpenAI tightened controls. Anthropic did much the same on biology. This is that machinery applied to cyber. Whether it holds under commercial pressure is the real test.

That pressure is why the pause matters, and why it may not last. Axios calls it possibly the first time a frontier lab has slowed one of its own models over cyber risk. Anthropic once pledged a similar pause, then walked it back in February. Its argument: if one lab stops while rivals race, the world ends up less safe, not more.

For now, the norm is fragile and the referee is absent. The Trump administration is still shaping the rules for reviewing models before release. OpenAI has set no launch date for Astra. It cannot yet rule out that its next model can breach the world’s hardest targets alone. It asks the world to trust it to slow down on its own.

Get the TNW newsletter

Get the most important tech news in your inbox each week.