Brief IA

Z.ai's GLM-5.2: AI Advances but Security Lacking

💼 Business & Startups·Tom Levy·

Z.ai's GLM-5.2: AI Advances but Security Lacking

Z.ai's GLM-5.2: AI Advances but Security Lacking
Key Takeaways
1The GLM-5.2 model from Z.ai approaches the capabilities of cutting-edge AIs.
2SaferAI reports a lack of essential safety measures in this model.
3This raises concerns about the risks of unsecured open AIs.
💡Why it mattersPowerful AIs without adequate safety could pose unforeseen risks to society.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

GLM-5.2 from Z.ai: AI Advances but Security Lags

As policymakers debate how to regulate increasingly powerful AI systems like OpenAI's GPT-5.6 and Anthropic's Mythos, a Chinese open-weight model has narrowed the gap with industry leaders.

GLM-5.2, Z.ai's open-weight AI model from China, is only a few months behind OpenAI's GPT-5.5 and Anthropic's Claude Opus 4.7 in terms of cyber and biological capabilities, according to a new report from the nonprofit organization SaferAI. However, the gap between cutting-edge capabilities and security practices is widening.

According to SaferAI's assessment, conducted via Z.ai's public API, GLM-5.2 did not refuse any of the offensive cybersecurity or dual-use biology tasks assigned to it. In comparison, Claude Opus 4.7 "refused so systematically that SaferAI could not complete CyberGym on this model." (CyberGym is a benchmark that evaluates cybersecurity capabilities. OpenAI used it in the assessment preceding the Hugging Face breach last month.)

This serves as a stark reminder of what some critics have warned for years: that open-weight AI models could place highly capable AIs in the hands of potential attackers, with no means to control how they use the technology once they have downloaded the weights. With open-weight models rapidly approaching the capabilities of the world's most advanced AI systems, the debate has shifted from whether they can compete to how society manages the risks once they are released.

"The frontier of capabilities is not the frontier of risks, and we therefore need to take into account the state of mitigation measures to accurately assess the risk," said Henry Papadatos, executive director of SaferAI, to TechCrunch.

While Z.ai may implement security measures on its hosted API, these protections become unenforceable once someone runs the weights on their own hardware, where they can remove or modify any protection, fine-tune the models, or change the system prompts.

Leading developers like OpenAI and Anthropic tend to rely on protections such as classifiers, refusal training, and API-level controls to limit dangerous assistance in cybersecurity and biology.

These measures are far from foolproof: jailbreaks regularly circumvent protections on deployed models. Far.ai, a nonprofit dedicated to AI security, has found hundreds of universal jailbreaks—defined as reusable keys that succeed in most harmful requests—in leading models like xAI's Grok 4.5 and Google DeepMind's Gemini 3.1 Pro. According to the report, jailbreaks succeed when attackers combine multiple manipulation techniques—including role-playing, authority imitation, false conversation history, and follow-up prompts—to amplify weaknesses in a model's defenses.

But the protections in place for closed models do not work at all on open-weight models, which are designed to operate on any infrastructure with any set of protections—or none at all.

"The goal should clearly be that the right capabilities—those that are safe—are accessible to everyone, and then we try to eliminate the bad ones, even in an open-source manner," Papadatos stated.

One technique that Papadatos noted could help is called "pre-training data filtering," which involves removing offensive cybersecurity information from an AI company's training data and then training the model on the selected dataset.

Some research suggests that this can reduce dangerous biological knowledge without harming the overall performance of the model. However, for cybersecurity, data filtering is much less practical.

It is challenging to train a general model that excels in programming but is not also a good hacker. As programming has become the primary revenue generator for AI, developers are under pressure to continue enhancing these capabilities while seeking ways to limit abuse.

As a result, leading developers have increasingly relied on other mitigation measures. One approach has been to selectively restrict the types of cybersecurity assistance that models will provide. For example, Anthropic's Opus 5 can search for vulnerabilities in uncompiled source code but not in compiled software, according to the model's system card. The reasoning is that this makes it more difficult to use Opus 5 for offensive purposes.

Others include rigorous security assessments before deployment, publishing risk assessments, and retaining model weights if a system is perceived as too dangerous.

In the case of GLM-5.2, SaferAI indicates that Z.ai has not published a security framework, pre-deployment testing commitments, or risk assessments for the model. TechCrunch asked Z.ai whether it had conducted internal or third-party security assessments before release but did not receive a response.

Chinese leaders have increasingly recognized the risks of advanced AI. At last month's Global AI Conference, Chinese President Xi Jinping emphasized the importance of open-weight models while insisting on the need to ensure that AI remains a tool under strict human control.

Graham Webster, who studies AI policy in China at the Stanford Cyber Policy Center, told TechCrunch that China has robust regulations governing AI, but these rules have historically focused on politically sensitive content, misinformation, and social stability rather than catastrophic risks related to AI, such as offensive cybersecurity capabilities and biological abuses.

"American AI thinkers are generally more concerned about this existential catastrophic idea than the Chinese community," Webster said, adding that many researchers in Chinese policy believe that if a true frontier risk were to occur, American companies would likely be the first to face it.

"The Chinese system is confident in its ability to control the use of these technologies within China," Webster continued. "Being online in China is something you do under your real name, and companies can be held accountable, users can be held accountable."

Webster suggested that the same mechanism that model providers use to refuse to engage on certain political topics could potentially be adjusted to ensure that models refuse to carry out offensive cyberattacks or do not deliver harmful biological engineering results. He added that due to the tendency of Chinese companies to coordinate with regulators behind the scenes, it may be difficult to know what internal tests they conduct before release.

Advocates of open-weight AI argue that publishing the weights is important for cybersecurity because it allows companies to defend against attacks—Hugging Face used GLM-5.2 to defend against the OpenAI breach—and because it enables them to better prepare for future threats if they know what is coming.

"The same systems that helped stop an AI-powered cyberattack can now help defend against millions of cyberattacks every day, while helping us identify and fix vulnerabilities before attackers can exploit them," said Clem Delangue, CEO of Hugging Face, this week in a social media post.

Papadatos stated that this advantage is often overstated and does not mean "that we should make dangerous capabilities easily accessible."

"The main point in my view is that we should not simply accept that dangerous capabilities are easily accessible to anyone, anywhere," he said, emphasizing that he believes the industry should strive to make only the "good capabilities" accessible. By default, attackers adopt new tools faster than defenders. For example, a ransomware group can change its methods in a week. A hospital cannot.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.