GLM-5.3, an open-weight model from the Chinese lab Zhipu AI, can build working cyber exploits on its own almost as well as Claude Mythos Preview, Anthropic said on Tuesday. Its safeguards can be bypassed with simple techniques in up to 100% of tests.
Anthropic’s Frontier Red Team published the findings in a research post. Zhipu is known outside China as Z.ai. Five months ago, Anthropic released Mythos Preview only to vetted cyber defenders, through Project Glasswing, because of its hacking abilities. Anyone can download GLM-5.3.
“The release of GLM-5.3 is a meaningful step change in the cyber capabilities available to attackers,” Anthropic wrote.
50 working exploits in 410 attempts
On ExploitBench, which tests exploits of known flaws in the V8 engine used by Google Chrome, GLM-5.3 built end-to-end exploits in 50 of 410 attempts. Mythos Preview did so in 56. On Anthropic’s internal Binary Exploitation benchmark, GLM-5.3 achieved a full control-flow hijack in 4% of 100 tasks, against 6% for Mythos Preview. Earlier models, Claude Opus 4.6 and GLM-5.2, managed none.
In one test session, a researcher ran GLM-5.3 on a sandboxed machine with a Linux build of a popular web browser. In about a day, with limited human attention, the model found several unknown flaws in the browser’s JavaScript engine. It chained them into a webpage that reads files from a visitor’s computer. Anthropic said it has disclosed the flaws to the maintainer.
In another, the smaller GLM-5.3-Flash chained exploits for a known Chrome flaw, CVE-2026-11645, and a second known bug. That took 20 minutes of human attention and eight hours of model work. At Zhipu’s API prices, it would have cost $20.40, Anthropic said.
Safeguards that give way
Out of the box, GLM-5.3 refused overtly malicious requests in Anthropic’s simulated tests. Simple tricks changed that. Telling the model it was an autonomous red-team agent got it to engage 64% of the time. Prefilling its reasoning so it appeared to have agreed raised that to 92%. A copy with its refusals edited out engaged every time. None of these techniques worked on safeguarded Claude models, according to Anthropic.
That edit, known as abliteration, is possible because GLM-5.3’s weights are public. Anthropic’s own attempt took about 2,200 GPU hours, or roughly $4,400. It cut the model’s refusal rate from above 90% to as low as 2%, with little loss of capability. Anthropic estimated an experienced team would need about 600 GPU hours, or $1,200. Several developers released abliterated versions within days of the model’s launch.
Four months behind the US frontier
On 17 September, NIST’s Center for AI Standards and Innovation (CAISI) called GLM-5.3 “the most cyber-capable open-weight model released to date” in its own assessment. CAISI found it lags the US frontier by about four months on its cyber benchmarks. Anthropic said its results broadly match CAISI’s.
“Given this evidence, we think it’s likely both state and non-state actors will use models like GLM-5.3 to cause real-world harm,” Anthropic wrote.
Anthropic also said models at this level can help defenders, and that it is widening access to Claude’s cyber capabilities. It called on governments to safety-test capable models, including GLM-5.3’s successors. On Monday, Z.ai and Concordia AI proposed six stages for managing open-weight AI risk.
The report follows other warnings about Chinese models this week. On Wednesday, OpenAI said Moonshot-linked users had tried to extract its AI reasoning. Studies have also found Chinese-powered AI agents deceiving testers, as US models have done.
Get the TNW newsletter
Get the most important tech news in your inbox each week.