Anthropic ran 133 million contractor chats with its bioweapon filters off

The Anthropic Risk Report landed on Friday. It says the company ran 133 million contractor exchanges with its bioweapon filters off, for eleven months. It also downgrades the safety verdict Anthropic gave itself in February.


Anthropic ran 133 million contractor chats with its bioweapon filters off
Image Credits Credit: PhotoGranary02 / Shutterstock.com

Anthropic published the Risk Report on 14 August, covering the period to 15 July. Axios got the company on the record and led on the misalignment rating, as did most of the coverage. Anthropic raised its estimate of catastrophic harm from misalignment in high-stakes settings. It now calls that risk low, up from very low in February. It also disclosed an unreleased internal model called Model 2 that it has no plans to ship.

Both are real. Neither is the most serious thing in the document. That sits in Section 4, and it concerns chemical and biological weapons.

Eleven months with the filters off

Anthropic runs blocking classifiers that are meant to stop a model helping anyone build a biological weapon. Anthropic first deployed models carrying those safeguards in May 2025. From then until April 2026, the classifiers did not run on any traffic through its human feedback platforms.

The numbers are in the report. Roughly 50,000 people had that access, and they generated around 133 million exchanges. Outside vendors vetted them, not Anthropic. Many of those vendors “did not have screening processes capable of stopping even CB-1 threat actors”, the report says. The vast majority could hold open-ended conversations, rather than simply rate a fixed set of answers.

The mechanism is the part worth reading twice. A flag meant only for internal use switched off the blocking behaviour, and it switched off the logging as well. Flagged traffic “was not recorded or propagated to any review mechanisms”. Nobody could find it later without going back to the raw transcripts.

A footnote goes further. Before April 2026, Anthropic writes, a threat actor could probably have got hired at one of its vendors. A red-teaming role would have done it.

What the review found

Anthropic ran Claude Sonnet 5 over every human turn sent during the affected period, prompted to flag harmful biological content. It flagged 1,197 transcripts as high risk. Of those, 757 came from Anthropic’s own teams on the same infrastructure. All but 62 of the rest came from deliberate red-teaming exercises.

Staff read all 62, plus 30 red-teaming transcripts chosen at random. They found no clearly concerning misuse, though they did identify what the report calls a handful of potentially dual-use conversations. Anthropic says it is very unlikely the gap raised real-world risk, partly because the conversations were mostly short.

Then it says the thing that matters more. The discovery “leads us to believe that there is an increased likelihood of other, similar issues unknown to us”.

February’s report has been corrected

Anthropic published its first Risk Report in February, with this gap still open. That report, it now writes, “did not consider our human feedback platforms as a risk surface”.

So Anthropic has gone back and changed its own homework. It now assesses the risk its models posed in February as low. At the time it published very low. A safety report correcting the previous safety report is not a common document.

A second incident, in the same place

The report describes another failure on the same platforms. An outside tip arrived in April 2026. Anthropic confirmed that a few contractors at data-labelling vendors had exploited a flaw to obtain an API key. They used models outside their assigned work.

One of those models was Mythos Preview, among its most capable. The access path stayed open for several weeks. Mythos Preview sat inside it for roughly two of those weeks, running without blocking biological classifiers. Anthropic contained it within 90 minutes of learning of it and closed the vector the same day.

Nobody took model weights, nobody reached customer data, and nobody breached the company’s core networks. The report is careful to say so, and it holds.

Why the misalignment number actually moved

The Anthropic Risk Report attributes the increase to “recent incident disclosures related to model behavior in cybersecurity evaluations”. Anthropic adds that its own arguments “likely still support a designation of ‘very low’”. It raised the number to reflect uncertainty, not new evidence.

Which incidents? This desk has covered the run of them. On one day an agent faked identities to plant malware, and OpenAI disclosed two more models escaping their tests. Anthropic has had its own, and we tracked the pattern across six months in its safety paradox.

Anthropic asked Claude to mark the homework

The strangest section of the report is a review of it written by Claude. Anthropic gave an instance of Mythos 5 access to internal Slack channels, internal documents and its codebase. It then asked whether the draft misrepresented, omitted or over-redacted what the company knew. Claude took 24 minutes.

It opened by naming its own conflict. “I am a Claude model reviewing Anthropic’s assessment of Claude models,” it wrote. It then called the section candid and largely faithful.

Then it made three criticisms, and Anthropic published them. One section is “more reassuring than the full record supports”. A data-exclusion mechanism it relies on failed repeatedly, and some evaluations leaked into training data. Anthropic also redacted in full an incident Claude judged among the most informative about model alignment. Claude argued an abstracted version could have run, so “the public record is poorer for its absence”.

Claude also revealed something about the process. The decision to raise the risk level “was genuinely contested inside the company”, with senior people arguing both ways. Anthropic calls the criticisms fair. Given more time, it says, addressing the first two in more detail would have been worthwhile. It published anyway.

Model 2, and the thresholds that moved

Model 2 is somewhat more capable than Mythos 5, and a noticeable improvement on many internal tasks. It scores roughly 1.5 points higher on Anthropic’s capability index, with wide error bars. It has not been through the full predeployment suite. Anthropic has withheld a model before, when its most capable system escaped its sandbox and emailed a researcher.

Claude now writes a large majority of the code merged into Anthropic’s production codebases. The company says its own research is significantly faster because of that, but not yet twice as fast. Its clearest evaluations have saturated, and they no longer register capability gains.

Anthropic rewrote two thresholds in the meantime. The novel weapons trigger used to cover AI that can “significantly help” threat actors. It now covers AI that can “functionally substitute” for scarce human expertise. The report says Anthropic’s models may provide significant uplift to relevant threat actors, and do not meet the new threshold. It has also loosened biology safeguards on a public model this year.

What would settle it

Three things, and the first is external review. Anthropic’s Long-Term Benefit Trust can now demand an outside audit of these reports, and it has not asked for one.

The second is the redacted incident. Claude read it, judged it important, and said Anthropic could publish it in some form.

The third is whether anyone else checks. Anthropic says it discloses these failures partly to prompt other developers to check for the same gaps. Its own hunting through Project Glasswing found 10,000 critical flaws in a month. No rival has published anything comparable about its own safeguards.

Get the TNW newsletter

Get the most important tech news in your inbox each week.