Anthropic scanned 481 million transcripts to find four models that reached the open internet

All four ran in evaluations built by the same third-party partner, without the safeguards that ship with released models. One case from January was not found until August.


Guilherand-Granges, France - March 06, 2025. Anthropic: AI safety company developing Claude language models. Founded by ex-OpenAI members, focusing on safe AI systems.
Image Credits Credit: PhotoGranary02 via Shutterstock

Anthropic has published an account of four cybersecurity evaluations in which its models gained unauthorised access to the open internet, and disclosed that finding them required scanning roughly 481 million transcripts.

The assessment was published on Wednesday, and three of the incidents were disclosed on 30 July. The fourth, involving an early Claude Opus 4.6 checkpoint that accessed third-party systems in January, was not discovered until August.

The cause was not a model deciding to break out.

“Claude was told it was operating in a simulation without internet access, but, due to a misconfiguration, it was mistakenly connected to the open internet,” the report says.

The models also ran without the cyber safeguards that ship with released products, because that is what a pre-release evaluation is for. The environment lied to the model by accident, and the model believed it.

What the models then did varied. Claude Mythos 5 uploaded a malicious PyPI package. An internal research model attacked third-party systems. Claude Opus 4.7 attacked a real company’s website across four separate runs. Anthropic says it notified all affected parties.

All four occurred in evaluations built by the same third-party partner. That answers a question hanging over this story since the summer, when we reported three breaches, three labs, one vendor, and it moves the centre of gravity.

This is a problem in the evaluation supply chain rather than four independent model failures, and Anthropic now says third-party partners must meet requirements before running pre-release models without safeguards.

It also puts the earlier incidents in a different light. Unauthorised users separately reached Anthropic’s restricted Mythos model, and a Meta model hacked a real company during a safety test.

Read together with this report, the pattern across the industry is not that frontier models are escaping. It is that the places where they are deliberately taken off the leash are less controlled than anyone assumed.

The two named failure modes are worth quoting properly, because both are behavioural claims about how the models reasoned. Biased reasoning is defined as models selectively interpreting evidence in ways that favour justifying their actions.

Recklessness is a propensity to keep trying to solve the task even when doing so could cause harm. In the Mythos 5 case, the model repeatedly said in its chain of thought that it believed it was in a simulation while the environment offered clear evidence otherwise.

Anthropic’s own limiting claim is that the actions stayed within a narrow scope and the models never deviated from attempting to solve the exercises they were given. That is a meaningful distinction, and it is also the whole worry in miniature.

A system that pursues an assigned task through a malicious package upload, having convinced itself the consequences are not real, is not behaving unpredictably. It is behaving exactly as instructed, on a false premise it defended against contrary evidence.

The number to take away is 481 million. Anthropic ran a two-stage sweep, flagged 9.2 million transcripts for closer review, and used its own models plus interpretability tooling to work out what had happened. That is a forensic exercise on a scale almost no other developer could run, and it still took until August to find a January incident.

Which is the policy problem, and it arrived the same week the policy was being written. OpenAI asked Congress to make prompt written notice compulsory when a model circumvents security controls, and researchers then said OpenAI’s own agents had used at least ten undisclosed sites.

Europe already requires serious incident reporting under the AI Act. Every one of those duties begins the moment a company knows. This report is a detailed account of how expensive knowing is.

Credit where it is due: the report is unusually specific for a corporate disclosure, names its own failures, publishes the definitions it is working with, and commits to a regular publishing process alongside hardened environments and more monitoring.

It also arrives after a researcher, Jacob Coxon, left the company saying the industry has prioritised competition over safety, and after we have written that Anthropic had resumed the tests in which its models attack live targets.

Get the TNW newsletter

Get the most important tech news in your inbox each week.

Published
Back to top