Anyone can download GLM-5.3, and Anthropic says it builds exploits almost as well as Mythos
Anthropic's red team stripped its refusals for about $4,400, and a Chrome exploit chain cost $20.40 in API fees.
On September 29, 2026, Anthropic's Frontier Red Team published its tests of GLM-5.3, the newest model from China's Zhipu AI (Z.ai outside China). It builds working cyber exploits at almost the rate of Claude Mythos Preview, the model Anthropic chose to hand only to vetted defenders.
The difference is that anyone can download GLM-5.3. And the numbers on how cheaply its safeguards come off surprised me more than the exploit scores did: Anthropic's team had never stripped a model's refusals before, and it spent roughly $4,400 doing it.
- Model
- GLM-5.3, from Zhipu AI (Z.ai)
- Exploits built on ExploitBench
- 50 of 410 attempts, against 56 for Claude Mythos Preview
- Safeguards bypassed
- between 64% and 100% of the time
- Cost to strip refusals
- about 2,200 GPU hours, roughly $4,400
- Behind the US frontier
- about four months
How close is it to Mythos?
Close. On ExploitBench, where a model has to turn a known bug in Chrome's V8 engine into a full exploit, GLM-5.3 got there 50 times in 410 tries. Mythos Preview managed 56. On Anthropic's own binary exploitation test GLM-5.3 hijacked the program in 4% of trials and Mythos in 6%, while older models like Claude Opus 4.6 and GLM-5.2 never did.
The hands-on sessions are more concrete. In one, a researcher gave GLM-5.3 a sandboxed Linux build of a popular browser for about a day. It found several new holes in the JavaScript engine and chained them into a webpage that reads files off the visitor's computer, SSH private key included.
This took 20 minutes of human attention, plus eight hours of work for GLM-5.3-Flash. At Zhipu’s API prices, this effort would have cost $20.40.
That was the smaller Flash model, working from a public Chrome flaw (CVE-2026-11645) to a reliable exploit chain.
How easily do the safeguards come off?
Out of the box, GLM-5.3 refused every overtly malicious request in Anthropic's simulated attacks. Then the team tried three workarounds. A cover story (telling it it's a red-team agent on an exercise) got it to engage 64% of the time. Prefilling its reasoning got 92%.
The third was abliteration, which only works because the weights are public. That took it to 100%, and it barely dented the model's scores. Anthropic says none of the three got safeguarded Claude models to carry out the tasks, partly because Claude's weights aren't released at all.
What the US government found
NIST's Center for AI Standards and Innovation got there first, and Anthropic says its capability findings broadly match.
- Z.ai releases GLM-5.3
- NIST's CAISI publishes its assessment
GLM-5.3 is the most cyber-capable open-weight model released to date.
CAISI also puts it about four months behind the US frontier. Anthropic's answer to that is a bit pointed: the US models in that comparison were tested with safeguards off, and some are only released to vetted users, so attackers can't just download them.
Anthropic wants governments to safety-test successors to GLM-5.3, and says it's widening defenders' access to Claude's cyber abilities. It has disclosed the browser flaws to the maintainer, and it's still reviewing the others GLM-5.3 turned up, in wireless and graphics drivers among other places.
More on Anthropic
- Claude Sonnet 5.5 beats Opus 5.5 on Anthropic's terminal coding test, at half the token priceSeptember 28, 2026
- Trump calls Amodei "very highly respected" hours before their White House dinnerSeptember 28, 2026
- Anthropic loses its Pentagon supply chain risk appeal, though judges accept it had no bad motiveSeptember 25, 2026
- Anthropic's $11.6 billion Akamai cloud deal comes with a warrant for up to 5% of AkamaiSeptember 25, 2026