GPT-5.5 and Mythos: Powerful AI, Growing Concerns
Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
GPT-5.5 and Mythos: Powerful AIs, Growing Concerns
Recent tests conducted by the AI Security Institute reveal that the GPT-5.5 artificial intelligence models from OpenAI and Mythos from Anthropic stand out for their capabilities in cyberattacks. The results of these tests raise increasing concerns.
This is the current issue with Mythos. This AI is impressively powerful, to the point that its creator Anthropic itself is calling for caution. Its deployment is already causing tensions, particularly from the White House, which fears uncontrolled use.
Comparable Performance in Cyberattack Tests
The tests show that GPT-5.5 and Mythos exhibit similar performance in complex cyberattack scenarios. In specialized benchmarks like CyberBench and the British simulation TLO in 32 steps, GPT-5.5 achieved a success rate of 71.4% on expert-level tasks. This score places it among the top-performing models at the moment.
Mythos is not far behind, with a success rate of 68.6% on the same tests. Although the gap is narrow, it is significant. Notably, GPT-5.5 managed to fully complete the TLO simulation in 2 out of 10 cases, while Mythos succeeded 3 times.
Advanced Hacking Capabilities
The skills of these AIs are no longer limited to technical assistance. They now execute complete attack chains, which is particularly concerning. The TLO simulation, for example, replicates a complex multi-step cyberattack, including reconnaissance, exploitation, privilege escalation, and lateral movement.
On the graph from the AI Security Institute, a clear trend emerges: as the tokens increase, the models progress through critical steps. GPT-5.5 follows a trajectory very close to that of Mythos, reaching advanced levels in areas like web exploitation or cryptographic analysis, typically reserved for human experts.
In detail, GPT-5.5 stands out for its consistency and stable progression through the steps, while Mythos shows sometimes faster but less consistent progress. Thus, GPT-5.5 becomes the second model capable of completing this simulation end-to-end, crossing a symbolic threshold in the field of cybersecurity.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.