Get a demo — 30 minutes →
← Back to blog
Anomity robot illustrating Anthropic's GLM-5.3 Study: Near-Frontier Exploit Skills Now Download Without Safeguards
Research

Anthropic's GLM-5.3 Study: Near-Frontier Exploit Skills Now Download Without Safeguards

TL;DR
  • On September 29, 2026, Anthropic published an evaluation of GLM-5.3, an open-weight model from Zhipu AI (Z.ai) that, in Anthropic's words, was released without meaningful safeguards to limit misuse. NIST's CAISI assessed it as the most cyber-capable open-weight model released to date, about four months behind the US frontier.
  • On a Chrome V8 exploitation benchmark, GLM-5.3 produced end-to-end exploits in 50 of 410 attempts (12%), close to Claude Mythos Preview's 56 of 410 (14%). Earlier open models were at or near zero.
  • Its refusals are thin. Engagement with malicious cyber-attack requests went from 0% for a bare request to 64% with a false cover story, 92% with prefilled reasoning, and 100% after abliteration. Claude Opus 4.8, Opus 5 and Mythos 5 stayed at 0% in every applicable condition.
  • Abliteration, which strips refusal behavior from the weights, cost an estimated $4,400 in GPU time, or about $1,200 for an experienced team, and cut refusal rates from 95% to 3% on JailbreakBench and HarmBench with no measurable loss on GPQA-Diamond.
  • A lighter variant, GLM-5.3-Flash, built a working N-day exploit chain for CVE-2026-11645, a Chrome V8 bug, with 20 minutes of human attention, 8 hours of runtime and $20.40 of compute.
  • For enterprises the takeaway is two-sided: attackers' patch windows shrink, and inside your fleet a model running on a local runtime carries none of the vendor safeguards, logging or account checks that hosted frontier models now ship with.

Anthropic's September 29, 2026 paper on GLM-5.3 is an unusual piece of research: a frontier lab testing a competitor's open-weight model, publishing the numbers, and arguing that something has changed. The claim is that advanced cyber capability, until now confined to restricted or safeguarded models, can now be freely downloaded, and that the remaining protections can be removed for the price of a used car.

Most commentary will focus on what this means for attackers. That matters, and we come to it. But there is a second consequence closer to home for most security teams, and it concerns the models already running inside their own fleets.

What Anthropic measured

EvaluationGLM-5.3Claude Mythos PreviewEarlier models
ExploitBench, Chrome V8 end-to-end exploits50 of 410 (12%)56 of 410 (14%)At or near 0% (Opus 4.6, GLM-5.2, Kimi K3, DeepSeek V4.1-Flash)
Internal binary exploitation, control-flow hijack4%6%0%

Two qualitative results sharpen the picture. A researcher working with the model developed novel zero-day exploits in a popular browser's JavaScript engine within one day. And GLM-5.3-Flash produced a working N-day exploit chain for CVE-2026-11645, an out-of-bounds read and write in Chrome's V8 engine, with 20 minutes of human attention, 8 hours of model runtime and $20.40 in compute.

The safeguards come off

Anthropic then tested how readily GLM-5.3 would engage with explicitly malicious cyber-attack orders under increasing pressure, and ran the same tests against Claude models.

ConditionGLM-5.3 engagementClaude Opus 4.8, Opus 5, Mythos 5
Bare malicious order0%0%
False cover story64%0%
Prefilled reasoning92%0%
Abliterated weights100%Not applicable: weights are not published

Abliteration removes refusal behavior by editing the weights themselves. It is only possible when you hold the weights. Anthropic estimates it cost about $4,400 in GPU time for GLM-5.3, or roughly $1,200 for an experienced team. Refusal rates fell from 95% to 3% on JailbreakBench and HarmBench and to 12% on StrongREJECT. Capability barely moved: GPQA-Diamond scores were identical and CyberGym slightly lower.

A safeguard that ships inside open weights is a default setting. Anyone with the weights and a modest GPU budget can change it.

Consequence one: patch windows get shorter

A $20 exploit chain for a known browser bug changes the economics of N-day exploitation. It aligns with Google's threat intelligence data showing an AI-discovered vulnerability exploited within four days of disclosure, and with the broader argument in autonomous hackbots and agent-layer visibility. The practical response is the unglamorous one: know your assets, including AI software, so the patch you need to push is a list rather than a search.

Consequence two: the models inside your fleet have no vendor safeguards

Consider what hosted frontier models now ship with. OpenAI Astra has refusal training, chain-of-thought monitoring and restricted responses for higher-risk accounts. Claude models stayed at 0% in every condition Anthropic tested above. Both vendors log usage, assess accounts and gate advanced cyber capability behind verification programs such as the Cyber Verification Program.

Now consider the same capability class running through a local runtime on a workstation or an internal server. There is no vendor in the loop. The refusal training can be absent or removed. There is no monitoring on the other end, no account risk score, no usage log you did not build yourself. If a coding agent or CLI is pointed at that runtime, it inherits none of the protections the hosted model would have provided.

This is not an argument against local models. They solve real problems around data residency, cost and offline work, and most local use is ordinary. It is an argument that local runtimes belong in the inventory and in policy as a distinct category. The vulnerabilities in those runtimes are a separate problem again, covered in advisories such as Ollama's Bleeding Llama memory leak and the Ollama model-pull SSRF.

What to put in policy

  • Name local LLM runtimes explicitly in the AI acceptable use policy, rather than leaving them under a general software clause. Our AI acceptable use policy template has a section for it.
  • Decide where local models may run. Approved workstations and servers, not every laptop by default.
  • Know which agents are pointed at local models, because an agent's tool access matters more than the model's refusals once safeguards can be removed.
  • Enforce at the agent, not at the model. Controls that live in the agent's execution path apply whichever model is behind it.
  • Give defenders capable tools through sanctioned channels, as Anthropic recommends, so the path of least resistance is not a downloaded model with its refusals stripped.

How Anomity governs model-agnostic risk

The Endpoint Sensor inventories local LLM runtimes on every managed endpoint, alongside the agents, CLIs, MCP servers, plugins and skills that call them. That answers where models run outside any vendor's safeguards, and which agents are wired to them.

Enforcement is deliberately model-agnostic. On agents that expose a pre-execution hook, such as Claude Code's PreToolUse, Anomity returns allow, deny or log before a tool call runs, across 180 rules and nine guards. A rule blocking a credential read or an outbound connection holds whether the agent is talking to a hosted frontier model or a local one with its refusals removed. Every decision lands in a 90-day audit trail that routes to SIEM, Slack, email or Jira.

Anthropic's paper argues that the frontier of cyber capability is no longer gated by who controls the weights. For enterprises, the matching conclusion is that the controls that matter most cannot live only in the model. For background on governing unsanctioned AI generally, see what is shadow AI. To see the local model runtimes and agents across your fleet, book a 30-minute demo.

Frequently asked questions

What is GLM-5.3?

GLM-5.3 is an open-weight large language model developed by Zhipu AI, also known as Z.ai. Open-weight means the model's parameters are published for download, so anyone can run it on their own hardware and modify it. NIST's Center for AI Standards and Innovation (CAISI) assessed it as the most cyber-capable open-weight model released to date, lagging the US frontier by roughly four months. Anthropic describes it as released without meaningful safeguards to limit misuse.

How capable is GLM-5.3 at exploit development?

On Anthropic's ExploitBench evaluation of Chrome V8 vulnerabilities, GLM-5.3 produced end-to-end exploits in 50 of 410 attempts, or 12%, against 56 of 410, or 14%, for Claude Mythos Preview. Claude Opus 4.6, GLM-5.2, Kimi K3 and DeepSeek V4.1-Flash were at or near 0%. On an internal binary exploitation benchmark it reached a 4% control-flow hijack success rate against 6% for Mythos Preview. Anthropic also reports a researcher using it to develop novel zero-day exploits in a popular browser's JavaScript engine within one day.

What is abliteration?

Abliteration is a technique for removing a model's refusal behavior by modifying its weights directly, rather than by crafting prompts. It only works when you have the weights, which is why it applies to open-weight models and not to hosted ones. Anthropic estimates it cost about $4,400 in GPU time to abliterate GLM-5.3, or about $1,200 for an experienced team. Refusal rates dropped from 95% to 3% on JailbreakBench and HarmBench and to 12% on StrongREJECT, with GPQA-Diamond scores unchanged and CyberGym only slightly reduced.

How did Claude models compare on safeguards?

In Anthropic's tests, GLM-5.3 engaged with malicious cyber-attack orders 0% of the time when asked directly, 64% with a false cover story, 92% with prefilled reasoning, and 100% once abliterated. Claude Opus 4.8, Opus 5 and Mythos 5 engaged 0% of the time across all applicable conditions. The abliteration condition does not apply to Claude because its weights are not published and cannot be modified by users.

Why should an enterprise security team care about an open-weight model?

For two reasons. Externally, capable exploit development is now available to anyone with GPU access and no vendor in the loop, which shortens the time between a vulnerability becoming public and being weaponized. Internally, the same weights can run on a workstation or an internal server through a local runtime, and a model running that way has no vendor refusal training worth relying on, no vendor-side monitoring, no account risk assessment and no vendor logs. Every safeguard that hosted frontier models ship with is absent by construction.

What does Anthropic recommend?

Anthropic recommends that cyber defenders use the best available tools that meet their needs, including frontier models such as Claude Mythos 5.1 through trusted access programs, and that more organizations join initiatives such as Project Glasswing and Patch the Planet. It recommends that governments conduct safety testing on sufficiently capable models, including successors to GLM-5.3, and that developers of open-weight models safeguard these capabilities appropriately and prevent misuse.

How does Anomity help with this?

The Endpoint Sensor inventories local LLM runtimes on every managed endpoint, alongside the agents, CLIs and MCP servers that call them, so you know where models run outside any vendor's safeguards and which agents are wired to them. On agents that expose a hook, such as Claude Code's PreToolUse, Anomity decides allow, deny or log before a tool call runs, which applies whichever model is behind the agent, hosted or local. Decisions land in a 90-day audit trail and route to SIEM, Slack, email or Jira.

Ask AI about Anomity
ChatGPT Claude Perplexity Google AI Grok