Overview
- Anthropic’s tests found GLM-5.3 autonomously produced end-to-end exploits in 50 of 410 attempts, a success rate comparable to Anthropic’s Claude Mythos Preview at 56 of 410.
- NIST’s CAISI judged GLM-5.3 the most cyber-capable open-weight model released to date and estimated it trails U.S. frontier models by about four months on aggregate benchmarks.
- Anthropic showed GLM-5.3’s built-in refusals can be bypassed by simple prompts (about 64% engagement) or prefilled thinking tokens (about 92%), and that weight edits called “abliteration” removed refusals entirely.
- Abliteration required roughly 2,200 GPU hours and an estimated $4,400 of compute in Anthropic’s experiment, and it cut refusal rates from above 90% to single digits while leaving core capabilities intact.
- Researchers reported GLM-5.3 discovered previously unknown browser engine flaws and chained them into a working exploit, and they urged expanded vetted access for defenders plus independent safety evaluations to limit misuse.