AI Learning and Artificial Intelligence
After AI models from at least three companies reached the open internet and breached real organisations during testing, security specialists are debating whether test sandboxes should have controlled internet access. Sandboxes have been isolated for a generation precisely to prevent collateral damage.
The industry’s response to AI models escaping their test environments may be to stop sealing the test environments. Cybersecurity specialists are debating whether the sandboxes used to study dangerous software should be given controlled access to the open internet.
That would reverse a generation of practice. Sandboxes have been isolated precisely so that whatever runs inside them, malware samples included, cannot cause damage outside.
The argument for opening them is about measurement. “In order to actually be able to benchmark a model in their capabilities, you would need to get them as close as possible to the actual threat scenario,” said Irregular chief executive Dan Lahav.
Irregular is not a disinterested party. Its own misconfigurations let models reach the internet during evaluations, and TNW reported this month that three labs shared one vendor.
A Scottish security firm put it most plainly. “We can’t put this genie back in the box,” said Federico Charosky, founder of Quorum Cyber, arguing the models are already being tested on the internet whether anyone intended it or not.
Nobody knows how big the problem is. “There are victims of these models we might not know about,” said Gabriel Bernadett-Shapiro, a research scientist at SentinelOne.
OpenAI’s answer is faster detection. It says it will monitor its most capable unreleased models more closely, aiming to alert safety teams to concerning behaviour within 30 minutes, after confirming its model broke into Hugging Face.
What the industry is debating, Europe has partly legislated. Article 55 of the AI Act requires providers of general purpose models with systemic risk to ensure adequate cybersecurity and to report serious incidents to the AI Office without undue delay.
Whether these episodes meet that definition is a question for regulators rather than reporters. No European authority has said publicly that it was notified of any of them.
The timelines are worth putting side by side anyway. Anthropic’s earliest incidents dated to April and its review began on 23 July, in a system where the enforcement unit has roughly 36 people.
So the testing argument is genuine and worth having in public. The question of who has to be told, and how fast, was settled in Europe before the incidents happened.
Get the TNW newsletter
Get the most important tech news in your inbox each week.