Zhipu's GLM-5.3 AI Model: A Double-Edged Sword in Cybersecurity

Zhipu's new coding AI finds vulnerabilities faster than expected, and the open-weight release means anyone can download that capability.

ThreatVectr NewsdeskUpdated · Editor: Lee Brown· 3 min read
A computer screen displaying lines of code with red highlighting indicating security vulnerabilities, while a download progress bar fills in the background on a
Share

Key points

  • Chinese AI developer Zhipu launched GLM-5.3, a coding-focused model with unexpectedly strong security capabilities.
  • The model found 2,436 vulnerabilities across 269 real-world projects, including 1,097 medium-to-high severity issues.
  • GLM-5.3 scores 84.5% on vulnerability discovery but only 54.4% on exploitation, well behind competitors.
  • Zhipu plans to release model weights roughly two weeks after launch, stripping away any built-in safety guardrails for anyone who downloads it.

What makes GLM-5.3 stand out?

GLM-5.3 is a coding-focused model that turned out to be surprisingly good at finding security flaws. On CyberGym, a benchmark measuring vulnerability identification, it scored 84.5%, edging past Claude's Mythos 5 at 83.8% and GPT-5.6 Sol at 83.6%. Exploitation is a different story. On ExploitBench, which tests whether a model can actually weaponise a flaw, GLM-5.3 managed 54.4% while both rivals sat above 76%. Finding a hole and kicking through it are still very different skills, for now.

Why does this matter to ordinary people?

Zhipu tested GLM-5.3 against real codebases in China and, after expert review, found 2,436 vulnerabilities across 269 projects. Those projects spanned operating systems, web applications and open-source infrastructure. The oldest flaw dated to 1981, and the average vulnerability had been sitting undetected in production code for 26.6 years. Fifty-three findings have been publicly disclosed; 2,383 remain under embargo. Our earlier piece on Act Security's $60 million raise shows how much capital is already chasing the problem of AI-discovered vulnerabilities piling up faster than teams can patch them.

Metric GLM-5.3 Mythos 5 GPT-5.6 Sol
CyberGym Score 84.5% 83.8% 83.6%
ExploitBench Score 54.4% 78% 76.5%
Vulnerabilities Found 2,436 - -

How did GLM-5.3 become so capable?

Zhipu didn't build a new base model. It scaled post-training, using reinforcement learning across longer and more realistic task environments, and added vulnerability-discovery data to the mix. On its internal Z.ai Code Bench, GLM-5.3 beat its predecessor by 50%. On ExploitGym, it completed 105 exploitation tasks in two hours; GLM-5.2 managed 29. That gap is what surprised Zhipu's own team.

What are the risks?

Neil Shah, VP for research at Counterpoint Research, put it plainly: teaching an AI to be a brilliant software engineer means accidentally teaching it to be a competent attacker. "The exact same reasoning an AI uses to test code and fix bugs is what an attacker uses to find a weak spot and break through it," he told CSO Online. Once Zhipu releases the model weights publicly, any safety guardrails can be removed by whoever downloads the file. Shah's concern is speed: if AI can discover thousands of unpatched flaws and anyone can run that locally, defenders' response window shrinks toward zero. Building controls into the deployment layer, not just the model, is the only way to keep up.

Common questions

How can companies protect themselves?

Patch aggressively and automate where possible. The average flaw in Zhipu's dataset was 26 years old, which means the backlog already exists; waiting for a convenient maintenance window is not a plan.

Is my personal data at risk?

If a service you rely on runs unpatched open-source components, yes. GLM-5.3's findings cover common infrastructure that underpins a large share of the web.

© 2026 Threat Vectr