Brief IA

GPT-5.5 and Mythos: Powerful AI, Growing Concerns

🤖 Models & LLM·Tom Levy·

GPT-5.5 and Mythos: Powerful AI, Growing Concerns

GPT-5.5 and Mythos: Powerful AI, Growing Concerns
Key Takeaways
1GPT-5.5 and Mythos show similar performance on cyberattack tests, according to the AI Security Institute.
2On CyberBench and TLO, GPT-5.5 achieves a success rate of 71.4%, competing with Mythos at 68.6%.
3Both AIs execute complete attack chains, going beyond the simple role of a technical assistant.
💡Why it mattersThese AIs could transform cybersecurity, posing national security risks and requiring urgent regulation.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

GPT-5.5 and Mythos: Powerful AIs, Growing Concerns

Recent tests conducted by the AI Security Institute reveal that the GPT-5.5 artificial intelligence models from OpenAI and Mythos from Anthropic stand out for their capabilities in cyberattacks. The results of these tests raise increasing concerns.

This is the current issue with Mythos. This AI is impressively powerful, to the point that its creator Anthropic itself is calling for caution. Its deployment is already causing tensions, particularly from the White House, which fears uncontrolled use.

Comparable Performance in Cyberattack Tests

The tests show that GPT-5.5 and Mythos exhibit similar performance in complex cyberattack scenarios. In specialized benchmarks like CyberBench and the British simulation TLO in 32 steps, GPT-5.5 achieved a success rate of 71.4% on expert-level tasks. This score places it among the top-performing models at the moment.

Mythos is not far behind, with a success rate of 68.6% on the same tests. Although the gap is narrow, it is significant. Notably, GPT-5.5 managed to fully complete the TLO simulation in 2 out of 10 cases, while Mythos succeeded 3 times.

Advanced Hacking Capabilities

The skills of these AIs are no longer limited to technical assistance. They now execute complete attack chains, which is particularly concerning. The TLO simulation, for example, replicates a complex multi-step cyberattack, including reconnaissance, exploitation, privilege escalation, and lateral movement.

On the graph from the AI Security Institute, a clear trend emerges: as the tokens increase, the models progress through critical steps. GPT-5.5 follows a trajectory very close to that of Mythos, reaching advanced levels in areas like web exploitation or cryptographic analysis, typically reserved for human experts.

In detail, GPT-5.5 stands out for its consistency and stable progression through the steps, while Mythos shows sometimes faster but less consistent progress. Thus, GPT-5.5 becomes the second model capable of completing this simulation end-to-end, crossing a symbolic threshold in the field of cybersecurity.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.