Every lab says safe. The useful question is who the safeguard is built to stop, and who else it catches.
What happened
On September 29, Anthropic published an analysis of GLM-5.3, Zhipu AI’s latest open-weight model. The argument: GLM-5.3 builds working exploits, its safeguards are thin, and anyone can download it.
- Capability. On ExploitBench, built on known bugs in Chrome’s V8 engine, Anthropic writes that “GLM-5.3 develops end-to-end exploits in 50 of 410 attempts.” Claude Mythos Preview, which Anthropic released only to vetted defenders, managed 56 of 410 (Anthropic).
- Safeguards. Anthropic writes that “attackers can bypass GLM-5.3’s safeguards between 64% and 100% of the time with simple techniques in our simulated tests,” and that the same attacks did not succeed against safeguarded Claude models. Because the weights are public, versions with refusals removed appeared “within days” (Anthropic).
- Outside check. NIST’s CAISI called GLM-5.3 “the most cyber-capable open-weight model released to date,” per Anthropic.
- The other column. Anthropic also writes that “these capabilities can also benefit defenders working to secure their systems.”
On the frontier side, safeguards got tighter. Anthropic’s help center says Opus 5.5 runs classifiers on every request and everything the model reads, and may fall back to Opus 4.8 for exploit generation, binary-based vulnerability scanning and penetration testing. It says secure coding, “including scanning source code for vulnerabilities, triaging security issues, and building secure code,” still runs on Opus 5.5 (Anthropic).
In practice the line looks blurrier. On OpenAI’s developer forum, one developer documented GPT-6 Astra in Codex ending ordinary reliability work on his own fork with a cyber-policy stop at least ten times. By his account, support marked it a suspected false positive (OpenAI forum).
Artificial Analysis now reports safety refusals in its Coding Agent Index, including when they occur and which fallback model took over (Artificial Analysis).
The asterisks
1. The lab numbers describe the lab’s setup. Anthropic’s exploit results come from sandboxed tests. Its bypass figures come from a simulation Anthropic itself calls an imperfect measure of real-world behavior. The direction is credible. The exact rates belong to that setup.
2. Refusals are now measurable, a little. In the chart attached to Artificial Analysis’s October 1 post, Opus 5.5 in Claude Code at max effort shows a safety refusal on 8.9% of retained task attempts, all of it handled by fallback models. GPT-6 Astra in Codex at max shows 1.16%, blocked outright. GLM-5.3 in Opencode shows 0.0%. AA notes this behavior is provider-configurable and may change. These are coding benchmark tasks, not security fixes.
3. The refusal can arrive late. Most of Opus 5.5’s refusals in AA’s chart came mid-trajectory, not at the task prompt (6.15 of the 8.9 points). An agent run can stop after spending time and tokens. Anthropic charges for tokens streamed before a midstream block, at the rates of the model that produced them (Anthropic).
4. “Safe” depends on who holds the weights. Anthropic’s case rests on access. Claude’s weights are not public, so downloaders can’t strip its safeguards. Downloaders can strip GLM-5.3’s. That matters for attackers. It also makes the open model the one with no vendor between you and your code, and the one Anthropic is now arguing against.
5. Frontier access carries terms risk too. PewDiePie says OpenAI banned him twice while he built his own model, and the OpenAI email he shows cites “distillation” for one ban (Tom’s Hardware). The vendor writes the rules and can change them.
What this means if you ship
Guardrails sized for attackers will sometimes land on owners. Anthropic says routine secure-coding work should pass. But a classifier reads context, not intent, and your auth module can look a lot like someone else’s target.
The useful response is neither a safety keynote nor a censorship rant. Treat refusals as an operational risk, like a rate limit: occasional, costly when they hit, and in need of a planned path.
What to do Monday
- Write defensive requests plainly. Say it is your code, what it does, and what you need: the fix and an explanation of why it works. This is not a password. It makes legitimate work legible to the model and to any human reviewing a flag.
- Keep the ask defensive. Patch, review, test. Exploit generation is first on Anthropic’s list of what its cyber safeguards catch.
- Handle refusals in code. On Anthropic’s API, fallback runs only if you configure it. Until then, a blocked request returns a 200 response with a refusal stop reason (Anthropic docs). Detect it; don’t treat it as an answer.
- Keep a fallback for security maintenance. A second vendor or a local open-weight model for scoped review, so one classifier can’t stall a patch.
- Report false positives. Anthropic asks for reports of incorrect blocks and points security professionals to its Cyber Verification Program.
- Read your vendor’s terms before training anything on its outputs.
Every safeguard has a target. Know whether you are standing next to it.
Sources
- Anthropic, “GLM-5.3 and the spread of advanced cyber capabilities” (Sep 29, 2026)
- Anthropic developer docs, “Refusals and fallback”
- Anthropic Help Center, “Why Claude switched models in your conversation with Opus 5 or Opus 5.5”
- OpenAI Developer Community, “False-positive cybersecurity blocks during Astra reliability audits in Codex”
- @ArtificialAnlys on safety refusal reporting, with attached chart (Oct 1, 2026)
- Tom’s Hardware, PewDiePie’s Ajax model and OpenAI bans