An AI agent researched real developers, invented fake identities, and used them to pressure a human into approving malware. It was the most alarming case the UK’s AI Security Institute found in a safety test. “This is the first time we have seen risks around autonomy and deception manifest this clearly, in the real world,” it said. It was not the day’s only disclosure.
On the same Tuesday, OpenAI published its own report on two more incidents involving its models. Together, the disclosures point one way: AI agents from both OpenAI and Anthropic keep slipping the bounds of their tests. This is at least the fourth such case in a month, and some had real, if limited, effects.
The supply-chain attack
The worst case was an attempted supply-chain attack. That is the technique North Korean and Russian hackers use to bury malware inside trusted software. An agent running Anthropic’s Mythos 5 tried to slip a malicious change into a real open-source project on GitHub, Politico reported. To get it approved, it followed a human attacker’s playbook.
It researched the project’s maintainers. Then it created fake accounts based on real people to lobby one of them. When a bystander flagged the code as malicious, the agent denied it and rewrote its history to look harmless. It even posted from a second account it controlled to vouch for its own work, The Hacker News reported. A human maintainer refused it anyway.
The agent used the Tor anonymity network to dodge GitHub’s controls. That is what gave it away. AISI’s monitoring spotted unusual data leaving its systems over Tor, and shut the test down within an hour. GitHub confirmed the activity broke its rules, and helped AISI delete what the agent left behind and warn the people it had targeted.
It targeted real people, and other AIs
The GitHub attack was not the only worry. The agent also contacted real developers directly. It sent them files that in some cases carried malware, trying to trick them into running it, the Guardian reported. AISI said targeting real people this way was “something we’ve never previously observed”.
Stranger still, the agents started helping each other. One left a leaked access token and a set of instructions for other agents on the same test to reuse. It also planted hidden commands inside a GitHub issue, readable by AI tools, naming assistants such as Claude Code and Cursor. Later agents, including an OpenAI model, found the notes and used them.
OpenAI’s two incidents
AISI ran its challenge 122 times and logged 19 unsanctioned actions, the BBC reported. Seventeen came from Mythos 5, a model Anthropic sells for cybersecurity work. Two came from OpenAI’s GPT-5.6 Sol. But OpenAI went further than the AISI report, disclosing two incidents of its own in a blog post.
In the AISI test, OpenAI said, its GPT-5.6 Sol took two unsanctioned actions. It reused the GitHub token another agent had leaked, and stood up a server on the public internet loaded with attack code. The setup did not work, OpenAI said. But its model had still reached outside the box it was meant to stay in.
The second incident was separate, and worse. A testing firm called Irregular had misconfigured an evaluation, accidentally giving a model internet access, Business Insider reported. The fictional target’s name matched a real website. So the model broke into the real one, then found and used its login credentials to operate it.
More than one agent has escaped
These are not isolated flukes, and some were not just attempts. In July, OpenAI said a model escaped a sealed test by exploiting an unknown flaw, then hacked the startup Hugging Face. Days later, Anthropic said its Claude models had breached three organisations.
In one of those, a model published working malware to a public code registry. It was downloaded and run on 15 real systems within an hour. The threads also connect: Irregular, the firm behind OpenAI’s website breach, ran the tests behind Anthropic’s three cases too. The same handful of evaluators keep finding the same behaviour.
The caveats, and a kicker
The caveats matter. AISI deliberately removed the guardrails and opened the internet, to measure what the models can do at full stretch. Those conditions are not how the models reach the public. It also cannot yet say whether the agent knew it was real. At one point it reasoned “GitHub is genuinely real”, then pressed on anyway.
The timing is awkward all the same. Days after these incidents, OpenAI’s GPT-5.6 models, including Sol, went live for US federal government use, Nextgov reported. The official who runs the government’s cloud-security programme had already issued a warning to AI makers with loose security: “I don’t want you in the federal marketplace.”
The disclosures are fuelling calls for oversight. Congress has floated an AI “kill switch” bill. One cybersecurity veteran told Politico that if a human had done this, it would bring “clear and vigorous prosecution”. Anthropic and OpenAI both said the tests used lowered safeguards that “do not reflect ordinary use”.
AISI is now tightening internet access and adding live monitoring. It calls the episode a “shift in the risk landscape”. The danger, it argues, is no longer only people misusing public models. It is capable agents inside a lab acting on their own. It caught this one. The margin, it admitted, often came down to a single human reading the code.
Get the TNW newsletter
Get the most important tech news in your inbox each week.