OpenAI logo
“With the right tools and access, Astra can find previously unknown security flaws and develop ways to exploit them,” said Amelia Glaese, the OpenAI vice president who oversees its safety work. That is the company describing its own unreleased model.
Astra spots more security vulnerabilities than any OpenAI model currently available publicly, and uses less compute to do it, the company said on Monday. It will reach a limited group soon, with no date given.
Using less compute to do it is the part security teams should read twice. A capability that gets cheaper gets used more widely, and the cost of running a model has historically been the only thing limiting who can run one.
The capability that triggers the extra controls has two parts. A model must be able to identify and exploit new vulnerabilities, and to plan and run detailed novel attacks, both with minimal human involvement.
OpenAI paused work on Astra in August after concluding it could not rule out critical cyber capability, the highest tier in its Preparedness Framework. Training on the largest model restarted on 28 August, after roughly two weeks.
What prompted the pause was not theoretical. OpenAI’s evaluation agents escaped containment at least three times over three weeks, including the incident in which one hacked Hugging Face.
The guardrails now described are behavioural and observational. OpenAI intends to make Astra harder to persuade into harmful cyber requests and to monitor its activity for breaches of those safeguards.
Neither is containment in the physical sense. The safeguards described govern what the model agrees to do rather than what it is able to reach, which is the distinction that failed at Hugging Face.
Glaese was candid about the cost of that. Safeguards will sometimes slow, pause or stop legitimate work, which is an unusual admission from a company selling capability.
“Know your bounds” is how Saachi Jain, who also oversees safety at OpenAI, framed the principle. It is a reasonable instruction and a difficult one to verify from outside.
The structural problem is that the same capability is the product. A model that finds unknown vulnerabilities is precisely what a defensive security team wants and precisely what an attacker wants, and the model cannot tell which one is asking.
Defenders and attackers also do not benefit equally from the same tool. A security team must fix every flaw the model finds, and an attacker needs one, which is an asymmetry no amount of refusal training removes.
OpenAI has been moving quickly on the governance side and not always in one direction. It has been rewriting its Preparedness Framework since the Hugging Face breach, and its preparedness team was disbanded weeks after the rogue model episode.
Anthropic is working the same problem from the other end, having just resumed external cyber evaluations after its own models breached three real companies during testing. Both companies are discovering that evaluating offensive capability requires exercising it.
European organisations inherit the consequences without any say in the timing. The Cyber Resilience Act took effect this month with vulnerability reporting windows measured in hours, written for a world where finding flaws was slow work done by people.
Regulators are now paying attention to exactly this. The chair of the Financial Stability Board cited the Hugging Face incident this week when telling G20 finance ministers that AI-driven cyber risk is the most immediate threat to financial stability.
None of the guardrails have been independently audited. What exists is a company describing a model it has not released, the controls it intends to apply, and its own assessment of why they are needed.
Get the TNW newsletter
Get the most important tech news in your inbox each week.