Anthropic: GLM-5.3 Reaches a Milestone in Binary Deployment

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
An internal test on 100 binary exploitation tasks highlights rare but real successes for GLM-5.3, achieving 4% of trials. Claude Mythos Preview performs better, at 6%, while earlier models remained at zero.
Previous generations with zero success, GLM-5.3 shows limited breakthroughs
Anthropic indicates that a significant threshold has been crossed in capabilities related to binary exploitation, as earlier models, including Claude Opus 4.6 and GLM-5.2, recorded no successes. In this area, GLM-5.3 achieves complete control flow hijacks in 4% of trials, while Claude Mythos Preview rises to 6% and surpasses GLM-5.3 in this evaluation. The measured criterion explicitly focuses on the ability to execute a complete control flow hijack, central to an examination of advanced cyber capabilities.
A protocol of 100 random tasks from an internal benchmark
The Anthropic Frontier Red Team tested several models, including GLM-5.3 and Claude Mythos Preview, on 100 tasks drawn from an internal benchmark of Binary Exploitation. The tasks were randomly selected, following a protocol presented by the team as an assessment of technical capabilities in the context of binary exploitation.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.