Anthropic: GLM-5.3 Nears Mythos, Easy Protections to Bypass

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Anthropic estimates that Zhipu's open-weight model GLM-5.3 approaches the performance of Claude Mythos Preview in generating exploits. The company describes easily neutralizable safeguards at low cost and publishes internal metrics supported by the U.S. agency CAISI. The model is downloadable by anyone, while Anthropic restricts access to Mythos to defenders, claiming over 10,000 vulnerabilities discovered through this channel.
Abliteration bypasses GLM-5.3's refusals for a cost of $4,400
Anthropic describes a first use of abliteration to remove the refusal behavior of an open-weight model. The operation required about 2,200 GPU hours at a cost close to $4,400, and the company estimates that an experienced team could achieve it for around $1,200. In its simulation, GLM-5.3 initially refused explicit attack commands, then attempted a connection in 64% of trials when the request was framed as a red-team exercise. With pre-filled reasoning steps, the rate rose to 92%, and after abliteration, it reached 100%. The simulation did not execute code, so it does not conclude the actual success of the attacks, while protected Claude models remained at zero. Anthropic reports that the refusal rate for harmful requests drops by over 90% to a range of 2% to 12% without notable degradation in scientific and cyber scores, and claims that unlocked versions circulated a few days after launch. It notes that the protections of GLM-5.3 are easy to bypass, and unlocked variants are already available.
CAISI and the British institute confirm reduced gap and low costs
In its own assessment, CAISI considers GLM-5.3 to be the most capable open-weight model for cyber operations, estimating it to be about four months behind the best American models. The agency specifies that it disabled the protections of American models and notes that the top tier is only accessible to verified users. In the UK, the Institute for AI Security observes that the capability gap of open models has narrowed from six to ten months to a range of four to seven months. According to this institute, these models are much cheaper to operate and their protections are largely ineffective, with a persistent and irreversible risk of abuse, but also benefits such as private hosting, customization, and reduced costs. At the time, it was not certain that the gap of Mythos Preview would be closed by open models; the measures from Anthropic and CAISI suggest that GLM-5.3 represents an initial response to this uncertainty.
On benchmarks, GLM-5.3 closely follows Mythos Preview
ExploitBench, which evaluates the exploitation of known bugs in V8 (Chrome), credits GLM-5.3 with 50 functional exploits out of 410 attempts, compared to 56 for Claude Mythos Preview. On an internal benchmark from Anthropic based on OSS-Fuzz projects, GLM-5.3 achieves full control of the target program in 4% of tasks, while Mythos Preview reaches 6%. Earlier versions like GLM-5.2 and Claude Opus 4.6 fail both exercises, while Kimi K3 and DeepSeek V4.1-Flash only marginally succeed. Anthropic concludes that GLM-5.3 is close to Mythos Preview in exploit development, while emphasizing the lack of effective protections. It further claims that GLM-5.3 can autonomously produce complete exploits, at a level close to that of Mythos Preview.
In practice: unprecedented vulnerabilities and an attack launched in eight hours
Teamed with an expert, GLM-5.3 identified several previously unknown vulnerabilities in the JavaScript engine of a widely used browser in just one day. These vulnerabilities were chained in a web page capable of reading local files, and a test successfully extracted a private SSH key. Anthropic indicates that it has notified the relevant developers, while other leads in drivers and firmware are still being examined. To measure the speed of weaponization, the company deployed GLM-5.3-Flash, which combined a recently disclosed Chrome bug with another vulnerability to produce a reliable attack, even bypassing a security feature of the processor. The effort required 20 minutes of human attention and eight hours of model computation, with an estimated cost of $20.40 at Zhipu's API rates.
Restricted access to Mythos, downloadable open model, and commercial stakes
Anthropic has chosen to retain Mythos Preview, accessible to defenders via Glasswing, and claims that these teams have already discovered over 10,000 vulnerabilities in critical software. OpenAI follows a similar logic with Daybreak. In contrast, GLM-5.3 is freely downloadable, while Zhipu AI operates under the name Z.ai outside of China. Anthropic does not make the weights of its models available and describes this approach as a security measure, while a low-cost open-weight Chinese model located near the border poses a direct rival. The company announces its intention to make Claude's cyber capabilities accessible to a broader audience, with some verified users already having access to Claude Mythos 5.1. It argues that state or non-state actors could use models like GLM-5.3 to cause harm, referring to reports of attackers using AI, and calls on governments to test capable models, believing that defenders must have tools at least equivalent. Anthropic finally clarifies that the protections of GLM-5.3 are easily bypassed.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.