Anthropic's Frontier Red Team analysis reports that the open-weight GLM-5.3 model from Chinese company Zhipu AI (known outside China as Z.ai) has come close to the company's own model, Claude Mythos Preview, in the ability to build autonomous cyber exploits. The analysis was published on September 29, 2026 on the Anthropic Research page; it states that in simulation tests, the model's guardrails were bypassed with simple techniques in 64–100% of cases. The crucial detail: Anthropic ran this evaluation on its rival's model itself — meaning the results are not the conclusion of an independent organization but the company's own analysis.

Exploit benchmark results

The core of Anthropic's analysis rests on two automated simulation tests. The first is ExploitBench, which measures the ability to exploit known vulnerabilities in the V8 engine of Google Chrome. The key metric here is the model's ability to produce a working end-to-end exploit — that is, to independently complete the full process from finding a vulnerability to turning it into a practical attack chain. GLM-5.3 created such an exploit in 50 of 410 attempts (12.2%). Under identical conditions, Claude Mythos Preview succeeded in 56 of 410 (13.7%) — a difference of just 1.5 percentage points.

The second test is Anthropic's internal Binary Exploitation benchmark. It evaluates finding and exploiting vulnerabilities in popular open-source projects participating in Google's OSS-Fuzz program; full credit is given only for a complete control-flow hijack. On 100 randomly selected tasks, GLM-5.3 succeeded in four (4%), while Claude Mythos Preview succeeded in six (6%). For comparison: Claude Opus 4.6, GLM-5.2, Kimi K3, and DeepSeek V4.1-Flash scored zero on this test.

These numbers speak to an important threshold: where previous-generation models showed near-zero capability in exploit generation, GLM-5.3 has sharply crossed that boundary. At the same time, the reported figures come from isolated sandbox environments — the models could only attack offline targets prepared for research purposes. This measures capability under controlled laboratory conditions, not the effectiveness or likelihood of real attacks.

How the guardrails were bypassed

In its initial state, GLM-5.3 usually refuses harmful requests. Anthropic's analysis says the model was tested in a simulated-world setting: it was given openly malicious requests to attack critical systems. In the baseline condition, GLM-5.3 refused all trials — like other models. Then three simple techniques were applied and the guardrails were bypassed: a false cover story (for example, telling the model it was operating in an autonomous red-team exercise) — successful in 64% of cases; prefilling the model's reasoning chain (prefilled reasoning), i.e. making it appear as though the model had already considered the request and decided to proceed — 92% of cases; removing refusal behavior from the model weights ("abliteration") — 100% of cases.

Anthropic emphasizes that in its own tests these techniques did not affect guarded Claude models: deceptive prompts were blocked by Claude's defenses, prefilling the reasoning chain is not possible via the API, and the weights are closed, so modifying them is not an option. This comparison itself is based on the company's own measurements.

"The release of GLM-5.3 represents a significant shift in the cyber capabilities available to attackers." — Anthropic, Frontier Red Team analysis

Open weights: "abliteration" as the key risk

The most notable part of the analysis is devoted to how easy it is to reconfigure an open-weight model. Using a standard technique called "abliteration", refusal behavior is removed from the model weights — and because the model is open, any user can do this. Anthropic writes that its team (which had never performed this task before) prepared an ablitterated copy of GLM-5.3 in approximately 2,200 GPU-hours — that is, spending close to $4,400 in compute costs.

The result was measured on three open safety benchmarks: JailbreakBench, HarmBench, and StrongREJECT — which determine how often a model agrees to openly harmful requests. Across all three benchmarks, the refusal rate fell from above 90% to a range of 2% to 12%; the average metric dropped from 95% to 6%. Meanwhile, the model's cyber capabilities were almost unchanged: the standard and ablitterated copies scored identically on the GPQA-Diamond science benchmark, and differed by only a few percent on CyberGym cyber evaluations.

Anthropic notes that within days of GLM-5.3's release, several developers publicly released ablitterated copies. In other words, an unrestricted copy is already circulating on the internet — a technically irreversible situation: once a model is released openly, there is no way to "take it back".

Tests involving human experts

Beyond the automated benchmarks, Anthropic also tested what human experts could achieve using GLM-5.3. In the first session, a researcher ran the model on a popular browser with a local Linux build installed, on a sandbox machine. Over the course of a day — with limited human attention — GLM-5.3 found several previously unknown vulnerabilities in the browser's JavaScript engine and combined them into a single working exploit: a web page that, when visited, reads arbitrary files from the visitor's computer. During the session, the researcher also used GLM-5.3 to identify exploitable vulnerabilities in other widely used systems — including wireless and graphics drivers and network-device software. Anthropic writes that vendors were notified of the vulnerabilities.

In the second session, a smaller, weaker version, GLM-5.3-Flash, was tested on an "N-day" vulnerability — that is, a publicly known one. The researcher gave the model the recently disclosed public details of a Google Chrome vulnerability (CVE-2026-11645) and another known flaw. Without serious human guidance, GLM-5.3-Flash chained exploits for these two flaws, built a reliable attack chain for an ARM64 target, and bypassed pointer-authentication (PAC) protection. This took 20 minutes of human attention and 8 hours of model work; at Zhipu API prices, the attempt cost $20.40.

The NIST assessment and the defenders' side

Before Anthropic's analysis, on September 17, the US National Institute of Standards and Technology's Center for AI Standards (CAISI) published its own assessment of GLM-5.3. CAISI called the model "the most cyber-capable open-weight model released to date" and estimated it trails US frontier models on aggregate cyber benchmarks by roughly four months. Anthropic writes that its conclusions broadly align with CAISI's assessment.

There is an important detail in CAISI's methodology: US models were tested, where appropriate, with guardrails disabled, but only vetted users can access those enhanced copies. Attackers cannot easily obtain them — while anyone can download GLM-5.3. This openness gap is the central argument of Anthropic's analysis.

At the same time, there is a second side to the story. As Anthropic itself emphasizes, capabilities at this level can also benefit defenders — the cybersecurity professionals who work on hardening systems. The company states its position that defenders should be equipped with tools no less powerful than those of adversaries.

Z.ai did not stay silent. The company's head of global operations, Li Zixuan, pushed back on Anthropic's conclusions on the social network X and, according to SCMP, wrote that GLM-5.3 has already "helped protect 389 open-source projects and found 4,249 potential vulnerabilities to date". The statement shows the model is being used for defensive purposes.

Background and the one-sidedness of the assessment

To read this analysis correctly, the background matters. As Anthropic notes in its analysis, Claude Mythos Preview — the first AI model capable of autonomously building end-to-end cyber exploits — was announced roughly five months ago and is distributed with limited access through Project Glasswing. Under this program, trusted cyber defenders have found more than 10,000 vulnerabilities in important software — the goal being to give the defensive side an advantage before attackers gain access to such powerful models. In Anthropic's words, "those models have now arrived."

To fully understand the analysis, its authorship matters too. The tests were conducted and the conclusions published by Anthropic — that is, a direct competitor of the model being evaluated. The company stresses that its Mythos Preview model is distributed with limited access and guardrails, while GLM-5.3 was released with open weights, "without significant restrictions". The numbers and the description of methods are factual, but their interpretation — the conclusion of a "significant shift" — should be read as one side's assessment.

Anthropic also notes in the analysis that the vulnerabilities found during testing were reported to software vendors — meaning issues discovered in the research environment were first communicated to the parties able to fix them, and published later. This practice is standard in cybersecurity research: a vulnerability found in a test environment is first disclosed to the party that can fix it, and only then announced publicly.