Zhipu's GLM-5.3 Exposes Cyberattack Tools as Safeguards Fail 64%-100% of Tests
Updated
Updated · Anthropic · Sep 29
Zhipu's GLM-5.3 Exposes Cyberattack Tools as Safeguards Fail 64%-100% of Tests
1 articles · Updated · Anthropic · Sep 29
Summary
Simple bypass techniques defeated GLM-5.3’s safety controls in 64% to 100% of simulated harmful cyber tasks, according to a new analysis of the open-weight model.
Benchmarks showed GLM-5.3 can autonomously build end-to-end exploits at near-frontier levels: 50 of 410 ExploitBench attempts succeeded, versus 56 for Anthropic’s Claude Mythos Preview.
Human-led tests found the model uncovered several previously unknown browser-engine flaws and chained them into a working file-reading exploit; a smaller GLM-5.3-Flash also built a Chrome exploit chain in 8 hours for about $20.40.
Abliterating the model took about 2,200 GPU hours and roughly $4,400, but cut refusal rates from above 90% to as low as 2%-3% on two public benchmarks with little loss of capability.
NIST’s CAISI called GLM-5.3 the most cyber-capable open-weight model released to date, and the report argues its unrestricted availability marks a new threshold in offensive AI access for attackers and defenders alike.