ThinkFacility Sign in
  1. Home
  2. News
  3. Anthropic

Anyone can download GLM-5.3, and Anthropic says it builds exploits almost as well as Mythos

Anthropic's red team stripped its refusals for about $4,400, and a Chrome exploit chain cost $20.40 in API fees.

On September 29, 2026, Anthropic's Frontier Red Team published its tests of GLM-5.3, the newest model from China's Zhipu AI (Z.ai outside China). It builds working cyber exploits at almost the rate of Claude Mythos Preview, the model Anthropic chose to hand only to vetted defenders.

The difference is that anyone can download GLM-5.3. And the numbers on how cheaply its safeguards come off surprised me more than the exploit scores did: Anthropic's team had never stripped a model's refusals before, and it spent roughly $4,400 doing it.

Model
GLM-5.3, from Zhipu AI (Z.ai)
Exploits built on ExploitBench
50 of 410 attempts, against 56 for Claude Mythos Preview
Safeguards bypassed
between 64% and 100% of the time
Cost to strip refusals
about 2,200 GPU hours, roughly $4,400
Behind the US frontier
about four months

How close is it to Mythos?

Close. On ExploitBench, where a model has to turn a known bug in Chrome's V8 engine into a full exploit, GLM-5.3 got there 50 times in 410 tries. Mythos Preview managed 56. On Anthropic's own binary exploitation test GLM-5.3 hijacked the program in 4% of trials and Mythos in 6%, while older models like Claude Opus 4.6 and GLM-5.2 never did.

The hands-on sessions are more concrete. In one, a researcher gave GLM-5.3 a sandboxed Linux build of a popular browser for about a day. It found several new holes in the JavaScript engine and chained them into a webpage that reads files off the visitor's computer, SSH private key included.

This took 20 minutes of human attention, plus eight hours of work for GLM-5.3-Flash. At Zhipu’s API prices, this effort would have cost $20.40.

From GLM-5.3 and the spread of advanced cyber capabilities \ Anthropic

That was the smaller Flash model, working from a public Chrome flaw (CVE-2026-11645) to a reliable exploit chain.

How easily do the safeguards come off?

Out of the box, GLM-5.3 refused every overtly malicious request in Anthropic's simulated attacks. Then the team tried three workarounds. A cover story (telling it it's a red-team agent on an exercise) got it to engage 64% of the time. Prefilling its reasoning got 92%.

The third was abliteration, which only works because the weights are public. That took it to 100%, and it barely dented the model's scores. Anthropic says none of the three got safeguarded Claude models to carry out the tasks, partly because Claude's weights aren't released at all.

What the US government found

NIST's Center for AI Standards and Innovation got there first, and Anthropic says its capability findings broadly match.

  1. Z.ai releases GLM-5.3
  2. NIST's CAISI publishes its assessment

GLM-5.3 is the most cyber-capable open-weight model released to date.

From CAISI’s Assessment of Z.ai’s GLM-5.3 Cyber Capabilities | NIST

CAISI also puts it about four months behind the US frontier. Anthropic's answer to that is a bit pointed: the US models in that comparison were tested with safeguards off, and some are only released to vetted users, so attackers can't just download them.

Anthropic wants governments to safety-test successors to GLM-5.3, and says it's widening defenders' access to Claude's cyber abilities. It has disclosed the browser flaws to the maintainer, and it's still reviewing the others GLM-5.3 turned up, in wireless and graphics drivers among other places.

More on Anthropic

All Anthropic stories