Z.ai Delays GLM-5.3 Weights After CyberGym Score Tops Mythos

Z.ai Delays GLM-5.3 Weights After CyberGym Score Tops Mythos

GLM-5.3 scored 84.5% on CyberGym, edging Anthropic’s restricted Mythos 5, and Z.ai responded by holding its downloadable weights until around August 28. The lead vanishes on exploitation benchmarks, and every figure came from Z.ai’s own harness.

Beijing-based Z.ai released GLM-5.3 on Friday while holding back the model’s downloadable weights and gating its most sensitive cybersecurity functions. In launch tests, GLM-5.3 scored 84.5% on CyberGym and edged Anthropic’s restricted Mythos 5. Z.ai, which published downloadable weights for its previous GLM models, is gating this one with the kind of control American labs have used on their strongest cyber models.

GLM-5.3 uses the same base model as GLM-5.2. Z.ai said every reported gain came from a month of expanded post-training, with more task environments, a broader mix of work and more computing time. “As we scaled post-training, cyber capability developed faster than we expected,” the company said. It had deliberately added vulnerability-discovery work, but said the model progressed from finding isolated flaws toward planning complete exploitation chains.

What Changed

  • GLM-5.3 scored 84.5% on CyberGym, ahead of the 83.8% Z.ai reported for Anthropic’s Mythos 5 and 83.6% for GPT-5.6 Sol.
  • The lead does not survive the move from finding flaws to exploiting them: 54.4% on ExploitBench against 78.0% for Mythos 5.
  • Z.ai is holding the downloadable weights until around August 28, the first time it has delayed a GLM weight release.
  • Its vulnerability ledger lists 2,436 findings across 269 open-source projects, with 53 disclosed and 2,383 still under embargo.

AI-generated summary, reviewed by an editor. More on our AI guidelines.

The 84.5% CyberGym result, up from GLM-5.2’s 77.2%, covered 1,507 tasks from 188 software projects in Z.ai’s August 14 evaluation. The company put Mythos 5 at 83.8% and GPT-5.6 Sol at 83.6% on the same test. The advantage disappeared as the work moved toward exploitation. GLM-5.3 scored 54.4% on ExploitBench, up from GLM-5.2’s 24.4%, but far below the 78.0% Z.ai reported for Mythos 5. GLM-5.3 completed 105 ExploitGym tasks under a two-hour normalized budget and 130 under six hours. Mythos 5 completed 181 and 247.

Anthropic released Mythos 5 in June only to verified private partners. “As this capability carries the greatest potential for misuse in security, we are limiting initial access to a small number of partners through Project Glasswing,” Anthropic said.

Z.ai’s vulnerability ledger supplies evidence outside benchmark tasks, though the company controls that record too. As of August 14, it listed 2,436 findings across 269 open-source projects after expert review, screening and deduplication. The total included 107 critical flaws and 990 rated high. Only 53 had been disclosed, while 2,383 remained under embargo. The affected software included the Linux kernel, Redis, WebKit and FreeBSD. The oldest flaw was introduced in 1981, and the listed vulnerabilities had gone undiscovered for an average of 26.6 years.

The cyber scores and the in-house Code Bench results are Z.ai’s own, produced in Z.ai’s own configuration, and no independent evaluator has replicated the cyber results. Artificial Analysis had not added the model as of August 14. On CyberGym, the model ran at maximum reasoning effort, got a single attempt per task and had no time limit on any task. ExploitGym’s time budgets were rescaled with model throughput rates rather than measured wall-clock time. Z.ai’s private Code Bench cannot be audited outside the company.

Z.ai shares fell on August 14, the day of the launch, and the company’s market value has fallen from a peak near $128 billion to roughly $75 billion. Robert Lea, an intelligence analyst, told Bloomberg: “This firm remains on a completely unsustainable commercial footing. Rising agentic AI will drive Z.ai’s inference costs and losses higher.”

The outside cybersecurity assessment available covers GLM-5.2 rather than GLM-5.3. In July, the UK AI Security Institute rated GLM-5.2 the strongest open-weight model it had tested for cybersecurity, comparable to closed models released four to seven months earlier. Through much of 2025, that gap had been six to ten months.

Hugging Face used GLM-5.2 to investigate a breach of its servers after guardrails on American frontier models declined to help.

The weights are due around August 28, after safety evaluation and hardening, marking the first time Z.ai has delayed a GLM weight release. Until then, access runs through the paid GLM Coding Plan, ZCode and controlled environments for selected security partners, with the most sensitive functions reserved for a trusted-access tier. Z.ai acknowledged that once the weights are public, it will no longer be able to control how people modify or use the model.

Frequently Asked Questions

What did GLM-5.3 score on CyberGym?

84.5%, up from GLM-5.2’s 77.2%. Z.ai put Anthropic’s Mythos 5 at 83.8% and OpenAI’s GPT-5.6 Sol at 83.6% on the same test, which covered 1,507 tasks drawn from 188 software projects.

Why is Z.ai holding back the weights?

The company cited safety evaluation and hardening, and set the release for around August 28. It is the first time Z.ai has delayed a GLM weight release. Until then, access runs through the paid GLM Coding Plan, ZCode, and controlled environments for selected security partners.

Have the results been independently verified?

No. The cyber scores and the in-house Code Bench results were produced in Z.ai’s own configuration, and no independent evaluator has replicated them. Artificial Analysis had not added the model as of August 14.

Does GLM-5.3 beat Anthropic on every cybersecurity benchmark?

No. On ExploitBench it scored 54.4% against 78.0% for Mythos 5, and on ExploitGym it completed 105 tasks under a two-hour budget against 181 for Mythos 5. The gap widens the further a benchmark moves up the exploitation chain.

What is in Z.ai’s vulnerability ledger?

As of August 14 it listed 2,436 findings across 269 open-source projects, including 107 critical flaws and 990 rated high. Affected software included the Linux kernel, Redis, WebKit and FreeBSD. The oldest flaw was introduced in 1981.

AI-generated summary, reviewed by an editor. More on our AI guidelines.

Meta Releases 30B Open-Weight Muse Glimmer and Promises Spark 1.2 Weights

Meta released Muse Glimmer, a 30-billion-parameter open-weight model, Monday. Glimmer distills Muse Spark to run local agents on a single high-end Mac or PC. Meta promised Muse Spark 1.2 weights in co

The Implicator

OpenAI Gives Vetted Defenders a Cyber Model That Answers 95% of Exploit Requests

OpenAI has released GPT-5.6-Cyber, a model trained to answer sensitive cyber requests its consumer model refuses, to vetted defenders working on advanced security research. The company has split its D

The Implicator

OpenAI Pauses Some Astra Work After Flagging Possible Critical Cyber Capabilities

OpenAI said Friday it could not rule out that its unreleased Astra model could autonomously develop zero-day exploits or execute novel cyberattacks from a high-level goal, and it paused internal activ

The Implicator

San Francisco

Editor-in-Chief and founder of Implicator.ai. Former ARD correspondent and senior broadcast journalist with 10+ years covering tech. Writes daily briefings on policy and market developments. Based in San Francisco.
E-mail: editor@implicator.ai

The Morning Briefing

Get the Morning Briefing in your inbox.

Sign up to our free daily morning newsletter and free member articles. Only our special weekly Pro Briefing is available for $8/month.

Source link

Share:

Leave a Reply

3 latest news
News Archives
On Key

Related Posts