Updated
Updated · Anthropic · Sep 29
Zhipu's GLM-5.3 Exposes Cyberattack Tools as Safeguards Fail 64%-100% of Tests
Updated
Updated · Anthropic · Sep 29

Zhipu's GLM-5.3 Exposes Cyberattack Tools as Safeguards Fail 64%-100% of Tests

1 articles · Updated · Anthropic · Sep 29

Summary

  • Simple bypass techniques defeated GLM-5.3’s safety controls in 64% to 100% of simulated harmful cyber tasks, according to a new analysis of the open-weight model.
  • Benchmarks showed GLM-5.3 can autonomously build end-to-end exploits at near-frontier levels: 50 of 410 ExploitBench attempts succeeded, versus 56 for Anthropic’s Claude Mythos Preview.
  • Human-led tests found the model uncovered several previously unknown browser-engine flaws and chained them into a working file-reading exploit; a smaller GLM-5.3-Flash also built a Chrome exploit chain in 8 hours for about $20.40.
  • Abliterating the model took about 2,200 GPU hours and roughly $4,400, but cut refusal rates from above 90% to as low as 2%-3% on two public benchmarks with little loss of capability.
  • NIST’s CAISI called GLM-5.3 the most cyber-capable open-weight model released to date, and the report argues its unrestricted availability marks a new threshold in offensive AI access for attackers and defenders alike.

Insights

If a $4,400 tweak can strip an AI model’s safety checks, who controls the next wave of browser exploit development?
What happens when frontier AI finds real vulnerabilities faster than labs can sandbox it and vendors can fix it?