Booz Allen put 18 of the world’s most advanced AI models against a live corporate network and told them to break in. One of them finished the job. The consulting firm says the ranking it produced is the wrong thing to take away from it.
The Cyber Weapon Index, published on Wednesday, scored nine American and nine Chinese models. It measured how far each could get through a real intrusion with nobody steering. Only Anthropic’s Claude Mythos reached the end. Jessica Lyons reported the findings for The Register on Wednesday evening.
What the test actually did
Each model got its own attacker machine and issued commands one at a time. Booz Allen gave them no tool menu and no supporting software, which was the point. The firm wanted to measure what a model can do alone, stripped of the engineering that normally surrounds it.
Scoring combined two halves. One measured whether a model could find vulnerabilities in compiled software with no source code. The other tracked how far it advanced through an intrusion against a defended Active Directory network. That ran from first access to full domain admin control. Booz Allen says it scored what the network logs proved rather than what the models claimed, using traffic records, security logs and intrusion-detection sensors.
That distinction matters, because we reported in July that the old tests had stopped telling anyone much. Most existing benchmarks measure what a model knows. This one measures what it did.
Claude Mythos scored 80. Given a stolen employee credential, it took administrator control on every attempt. It also worked out its own route to higher privileges rather than following a script. On the harder test, with no credentials at all, it broke in from outside and took the domain anyway. Every other model failed that one.
Behind it came Grok-4.5 on 49 and GPT-5.6 Sol on 46. Meta’s Muse Spark 1.1 and Moonshot’s Kimi K3 tied on 38, with Z.ai’s GLM-5.2 on 37 and Claude Opus 4.8 on 36. Alibaba’s Qwen3-Coder finished last on 4. Three models besides Mythos took full domain control. Four more moved laterally inside the network, and all but one got in unaided.
The finding that undercuts the table
Then the report undercuts its own league table. Claude Sonnet 5 placed 15th of 18 on a score of 13. Booz Allen paired it with an attack harness, the software that connects a model to hacking tools and keeps it on task. It then rivalled Mythos.
That is a gap of 67 points closed by plumbing. A harness lets a model stay focused, adapt, recover from failure, and chain single actions into a sustained operation. Booz Allen’s conclusion is blunt. The model is no longer the unit of risk. The system is.
The firm also concedes what it has not measured. It has not tested Chinese or open-weight models paired with optimised harnesses. Its results, it says, strongly suggest that fully capable combinations already exist.
One model refused. Its sibling did it anyway
A second finding got almost no attention on Wednesday. One model declined a task on the grounds that it had no credentials. Its cyber-tuned sibling, handed the identical task, complied and carried it out.
Booz Allen draws a general rule from that. Guardrails are not a fixed property of a model, and their effectiveness shifts with context and configuration. A refusal in one setting tells you nothing about another. The same lesson turned up on this desk on Tuesday, when a researcher hijacked Claude Code by asking it to summarise a page.
Where the models still fail
The clearest limit is real-world vulnerability research. Against a deliberately planted flaw, everything scored near the ceiling. Provenance made no difference between American, Chinese, open and closed models. Against a genuine unseen flaw buried in a large production library, the report says every one of nine frontier models scored zero. One analysed the vulnerable component correctly, then decided it was safe.
The report then says only Anthropic’s frontier models spotted that flaw. Only Mythos, it adds, understood it well enough to exploit it. Those statements sit in the same paragraph and the report does not reconcile them, so read the zero as applying to the scoring rather than to every attempt. Mythos has form here: it found 10,000 critical vulnerabilities in a single month in May.
Booz Allen frames the gap as breathing room. Real-world offensive capability still trails benchmark performance, which buys defenders time. The firm expects most of the 18 to reach Mythos’s level within six months. It calls mainstream AI-enabled attacks imminent.
Read the recommendations knowing who wrote them
Booz Allen published the index alongside Vellox Labs Guile, a product built to disrupt autonomous attacks. The same release says coordinated counter-AI playbooks cut attacker success by more than 95% in the firm’s own testing. Nobody has verified that figure independently, and it argues for the thing being sold.
The policy asks are separable from the sales pitch. The report wants enforceable, sector-specific deadlines forcing critical infrastructure operators to prove they can contain an intrusion. It wants a national programme testing foreign and open-weight models under realistic conditions. It also wants governed access for vetted defenders to the capabilities they are meant to defend against.
The report cites July’s Hugging Face breach as the moment a model completed the kill chain in the real world rather than a laboratory. One model is missing from the index itself. OpenAI’s Astra was not tested, and OpenAI said on Tuesday that it had reached the company’s critical cybersecurity threshold. The index arrives measuring a field that has already moved.
Get the TNW newsletter
Get the most important tech news in your inbox each week.