China’s AI startup Z.ai says open-source GLM-5.3 beats Anthropic’s Mythos 5 in cybersecurity test
Less than two months after China’s Z.ai surprised the AI industry with GLM-5.2, a model that beat OpenAI’s GPT-5.5 on key benchmarks and narrowed the gap with leading U.S. AI labs, the Chinese startup is back with another challenge to the frontier.
This time, it’s cybersecurity.
Z.ai said Friday that its new open-source GLM-5.3 model scored 84.5% on CyberGym, a benchmark that tests an AI model’s ability to inspect code, find software vulnerabilities, and verify that the flaws are real. That narrowly topped the 83.8% score Z.ai reported for Anthropic’s restricted-access Mythos 5 model.
China’s Z.ai Says GLM-5.3 Tops Anthropic’s Mythos 5 in AI Vulnerability Detection

GLM-5.3 vs Anthropic’s Mythos 5 Performance Evaluation (Credit: Z.ai)
The result doesn’t mean GLM-5.3 has overtaken Mythos 5 overall. Anthropic’s model still held a substantial lead in turning discovered vulnerabilities into working exploits. But coming on the heels of GLM-5.2’s strong showing against leading U.S. models, the latest results add to a trend that is becoming harder to dismiss: China’s top open-source AI models are closing the capability gap faster than many expected.
“We evaluate GLM-5.3 across three benchmarks covering different stages of vulnerability analysis and exploitation. On CyberGym, which starts from white-box source code and tests whether the model can identify and validate vulnerabilities by triggering faults, GLM-5.3 scores 84.5%, up from GLM-5.2’s 77.2% — the best result on the benchmark, ahead of Mythos 5 (83.8%) and GPT-5.6 Sol (83.6%),” Z.ai said on the GLM-5.3 release page.

Z.ai’s claims have not been independently verified.
The gap becomes clearer once the models move from finding vulnerabilities to exploiting them. On ExploitBench, GLM-5.3 scored 54.4%, compared with 78.0% for Mythos 5. In a separate timed test, Z.ai said GLM-5.3 completed 105 attack-development tasks in two hours and 130 in six hours. Mythos 5 completed 181 and 247 tasks over the same periods.
On ExploitBench, which requires deeper reasoning about real vulnerabilities and their exploitation, GLM-5.3 reaches 54.4%, more than doubling GLM-5.2’s 24.4%, while Mythos 5 and GPT-5.6 Sol score 78.0% and 76.5%, respectively. On ExploitGym, which measures how many exploitation tasks a model can complete under time-normalized budgets, GLM-5.3 completes 105 tasks within two hours and 130 within six hours, compared with 29 and 39 for GLM-5.2; budgets are normalized across models using per-model throughput figures, detailed in the footnotes. Mythos 5 remains well ahead at 181 and 247 tasks.
That distinction matters. Finding a security flaw and producing a working exploit are very different capabilities, and the latter carries far greater potential for misuse.
Anthropic has kept Mythos, a version of its Claude Fable 5 model with certain cybersecurity safeguards removed, behind restricted access for vetted organizations. The company plans to release GLM-5.3 publicly in about two weeks after completing security assessments and strengthening safeguards.
Its most sensitive cybersecurity capabilities will remain limited to verified users through what Z.ai calls a “trusted access” program. Initial access will go to selected launch partners before opening to a broader group.
China’s open-source AI push enters a more sensitive phase
The release strategy may prove as significant as the benchmark scores.
“To the best of my knowledge, this is the first time a Chinese lab is publicly justifying a delayed open release of model weights with safety considerations,” Gabriel Wagner, an AI governance researcher at Beijing-based consultancy Concordia AI, told Reuters in a statement.
“This shows that open-weight risk management practices in China are becoming more sophisticated.”
Z.ai said GLM-5.3 includes systems for screening risky requests, monitoring model activity and rejecting malicious tasks. The goal is to separate legitimate security work, such as bug fixing, cybersecurity education and authorized testing, from attempts to misuse the model.
That becomes much harder once model weights are publicly available. Developers can modify open models, remove safeguards, or connect them to external tools, leaving model makers with far less control over how their systems are used.
Z.ai is leaning into that tension rather than retreating from open source. The company announced an “Open Source Shield” initiative to audit selected open-source projects, provide model access for defensive security work, and add code-auditing features to its ZCode programming product.
“In this spirit, Z.ai appears to be proposing a kind of ‘Project Glasswing’ with Chinese characteristics that sees openness as an asset rather than a drawback,” Wagner said.
Another reason to watch Z.ai closely is this.
New York-based Hugging Face said last month that it used Z.ai’s previous GLM-5.2 model to help defend against a cyberattack by a rogue OpenAI agent that breached its systems. Chinese cybersecurity company 360 has separately claimed that its Tulongfeng vulnerability-discovery system reached capabilities comparable with Mythos, though those claims have not been independently verified.
GLM-5.3 takes a different approach. It is a general-purpose coding model rather than a dedicated cybersecurity system. Z.ai said it shares the same base model as GLM-5.2, with longer and more varied post-training and reinforcement-learning environments producing its stronger security capabilities.
That makes the progression from GLM-5.2 particularly interesting. Z.ai was already attracting overseas developers with coding and agent capabilities approaching leading U.S. models at lower prices. Now the company is pushing that same model family into one of AI’s most consequential applications.
The benchmark race is far from settled. Mythos remains substantially stronger at exploit development, and Z.ai’s GLM-5.3 results still need independent verification.
But the direction is getting harder to ignore. First coding and AI agents. Now cybersecurity. The competition between Chinese open models and America’s leading closed AI systems is moving into capabilities where the stakes are much higher than benchmark bragging rights.

