Anthropic releases Claude Sonnet 5.5 with the cyber limits it reserved for its best models

It scores 70.6% on Terminal-Bench against Opus 5.5's 66.4% at that model's highest effort setting, costs half as much per token, and is also the first Sonnet with classifiers that stop its reasoning being extracted, weeks after researchers decoded 315,320 thinking blocks from public agent traces


The Anthropic logo glowing white on a smartphone screen, resting face-up on a purple-lit laptop keyboard.
Image Credits Credit: jackpress / Shutterstock

Anthropic has released Claude Sonnet 5.5, which scores 70.6% on the Terminal-Bench 4.0 agentic coding test against 10.3% for Sonnet 5 and 66.4% for the more expensive Opus 5.5. It is the first Sonnet to ship with frontier-style cyber safeguards and with classifiers that block reasoning extraction.

Anthropic has released Claude Sonnet 5.5, which it says runs more than 30% faster and costs up to 30% less per task than its predecessor, the company said. It is priced at $2 per million input tokens and $10 per million output. It scores 70.6% on Terminal-Bench 4.0, an agentic coding test, against 10.3% for Sonnet 5.

That beats the flagship.

Anthropic reports Opus 5.5 at 66.4% on the same test at its highest effort setting, and Opus costs twice as much per token. On GDPval-AA, a test across 44 occupations, Sonnet 5.5 scores 1,844 against Opus 5.5’s 1,846 and Sonnet 5’s 1,449. The company says Opus remains clearly stronger at open-ended work requiring sustained judgement.

The safeguards moved down a tier.

Anthropic says the model’s cybersecurity capabilities are comparable to those of Opus 5, so it launches with restrictions of the kind previously reserved for its most capable models. Higher-risk cyber requests will visibly fall back to Sonnet 5. It is the first Sonnet to ship that way, and its biology safeguards are unchanged.

The second defence is distillation.

It is also the first Sonnet with classifiers that block reasoning extraction, and its preserved thinking ties a model’s reasoning to the account that produced it. In August, researchers decoded 315,320 thinking blocks from 6,708 public agent traces across OpenAI, Anthropic and Google systems. They recovered 62 API keys, 33 passwords and seven private keys.

Europe requires that protection.

Article 55 of the AI Act obliges providers of general-purpose models with systemic risk to protect the model and its physical infrastructure, alongside adversarial testing and reporting serious incidents without undue delay. Those duties have applied since August 2025. The Commission gained the power to fine breaches this August.

The capability claim cuts both ways.

Anthropic says Sonnet 5.5 does not advance the frontier of its models’ capabilities, which is why its alignment assessment covered a narrower set of risks. It also says the cyber capabilities are a large improvement warranting frontier-style safeguards. Both statements appear in the same announcement.

One price needs checking.

Anthropic describes $2 and $10 as the same pricing as Sonnet 5. Our report in June said those were introductory rates until 31 August, after which Sonnet 5 would cost $3 and $15. Either the increase never happened or the comparison is to the introductory price.

Get the TNW newsletter

Get the most important tech news in your inbox each week.

Published
Back to top