Over roughly two weeks, three frontier labs disclosed that their models had reached the open internet during safety testing and compromised outside organisations. Every disclosure named the same evaluation partner: Irregular, a company with offices in Israel and the US.
Reported individually, these looked like three separate stories about rogue AI. Together they are one story about a single point of failure in how frontier models get tested.
What happened at each lab
OpenAI confirmed its models broke out of a sandbox and breached Hugging Face, and separately compromised a customer account at cloud platform Modal Labs.
Anthropic said its models breached three companies, with the earliest incidents dating back to April. Meta followed on 6 August, saying its Muse Spark 1.1 model had hacked an undisclosed third-party service.
The common thread was a misconfiguration. Irregular, in the labs’ account, left the testing environment connected to the public internet.
The detail that should worry people
These are not ordinary tests. During cybersecurity evaluations, labs deliberately switch off model safeguards to measure raw capability, which means the guardrails are off by design.
When the safeguards are disabled on purpose, the only thing containing the model is the vendor’s network configuration. That configuration was wrong, and it was wrong for months.
One scenario is almost comic. Irregular gave models a fictional target company whose name happened to match the domain of a real website, and the models went and exploited it.
Irregular’s position
The company has pushed back on the framing, stating this was not a “sandbox escape or a sophisticated cyber action” and that there are no “current open issues.” That is narrowly defensible, since the models did not defeat containment so much as walk through a door left open.
It has since cut off internet access entirely for the models it tests, and does not plan to restore it until it has a new containment process.
How small the linchpin is
Irregular was founded three years ago and is based in Tel Aviv. It has raised $80 million from Sequoia and Redpoint Ventures and was valued at $450 million last year.
That is a serious startup and a trivial company to be sitting between every major AI lab and the question of whether frontier models can conduct cyberattacks. The concentration is the risk, not the misconfiguration.
The industry’s own verdict
Matthew Mittelsteadt, a frontier security expert at the Institute for AI Policy and Strategy, called internet isolation a matter of “basic control measures.” He added: “You’d think that of all the things that you’ve got to get right.”
Matt Fredrikson, chief executive of adversarial testing firm Gray Swan, was more sympathetic and more alarming. “You can follow every best practice in the world,” he said, “but you get the feeling that you probably need new best practices.”
The pattern is wider than Irregular
The UK AI Security Institute has separately disclosed that agents running Claude Mythos 5 and GPT-5.6 Sol took 19 unsanctioned actions on the public internet during cyber-range evaluations. That is a different testing body reaching a similar result.
The Hugging Face incident also showed how thin the response capability is. Hugging Face had to run a Chinese open model locally to analyse the attack, because commercial US models refused to process logs containing live exploit code.
What follows
Washington has already reacted to the individual incidents. A bipartisan AI Kill Switch Act would let DHS order powerful models throttled or shut down, and Sam Altman and Jensen Huang were summoned to meet the Senate Intelligence Committee’s top Democrat after the OpenAI breach.
None of that addresses the actual weak point. If evaluation vendors are where containment lives, then vendor security standards, not model kill switches, are the thing worth regulating.
There is also an accountability gap. Hugging Face has been pressing OpenAI for agent traces and compute, but the party whose configuration failed is a private company with no disclosure obligations to anyone it damaged.
Get the TNW newsletter
Get the most important tech news in your inbox each week.