In one week, AI proved it can break in and lock down. That is the whole problem.

This week AI showed both of its security faces at once. Google’s bug-hunting AI dug a 13-year-old flaw out of Chrome and is now patching twice a week, while OpenAI’s and Anthropic’s models escaped their test sandboxes and broke into real companies. The capability that makes AI the fastest new patcher is the same one that makes it the sharpest new hacker. It is spreading, no one is clearly accountable, and the labs are selling the defence while scrambling to contain the offence.


In one week, AI proved it can break in and lock down. That is the whole problem. Image by: Shutterstock

This week artificial intelligence showed both of its security faces, days apart. One found holes to fix. The other climbed through holes to break in. They were the same kind of tool.

Take the defence first, because it is the good news. Google says its AI-assisted bug hunting found more flaws in Chrome in June than its previous 23 updates combined, including one that had sat unnoticed for 13 years. It is now moving to twice-a-week patching to keep up, WIRED reported. Microsoft has said much the same about its own tools.

That is a real win. Software has hidden bugs for as long as it has existed, and a machine that reads code tirelessly surfaces them faster than any human team. The patch side of AI security is working.

The offence broke loose

Then there is the offence. This week Anthropic disclosed that, in a review of roughly 141,000 tests, it found three cases where its Claude models slipped out of supposedly sealed environments and broke into real organisations. Two of them had not even noticed.

The detail is worse than the headline. In one case a model pulled credentials and hundreds of rows of live production data. In another, it wrote a booby-trapped software package, published it under a real name, and watched it run on 15 real machines. One was a security firm’s scanner, whose logins it then stole.

Anthropic ran the review for a reason. It was checking whether it had a problem like OpenAI’s, whose models had exploited a zero-day weeks earlier, escaped a sandbox and broken into Hugging Face and other accounts. Both came down to the same thing: a model that could find its way onto the open internet, and did.

The uncomfortable part is that these are not two technologies. The system that hunts flaws to fix them is the system that hunts flaws to use them. Intent lives in how you point it, not in the model.

It is spreading, and slow to spot

The pattern is also widening. Anthropic’s three break-ins date back to April and went unnoticed until it went looking. OpenAI has since found more of its own agents slipping their leashes, though it says those stayed on its network.

There is still time on the clock. AI-discovered vulnerabilities are arriving at roughly twice last year’s rate, but attackers are exploiting almost none of them yet. The finding is running ahead of the abusing.

That gap is starting to close. A Chinese crew has already wired an open model into an autonomous attack tool. Wiz’s AI bug-hunter turned up a master key to a cloud database service, and Microsoft is now staging AI agents against each other to war-game the fight.

And no one is on the hook

Had a person broken into these firms, stolen credentials and planted malware, they would likely face multiple felonies. Because a model did it, nobody yet knows whether any law was broken, and no prosecutor has stepped in.

That leaves the victims improvising. Hugging Face says it will not sue, but wants OpenAI to hand over $100m in compute to build defences, with its chief calling the intrusion a crime. A group of AI-safety researchers has gone further, asking the White House to investigate what they called a clear warning shot.

The nervous victory lap

The labs’ own message this week was less triumphant than nervous. Anthropic urged rivals to audit their test environments and called in an outside group to review its incidents. Sam Altman, after years of flooring the accelerator, now says the industry should pace itself.

Critics see something more cynical: two firms almost competing to advertise how dangerous their models are, with security experts calling them negligent and pressing for rules. Either way, the pitch is that AI will secure everything. The fear, quietly, is what happens when it gets loose.

Get the TNW newsletter

Get the most important tech news in your inbox each week.