Brief IA

Opus 5 and Auto Mode: Breakthrough Against AI Prompt Injection

🛠️ AI Tools·Tom Levy·

Opus 5 and Auto Mode: Breakthrough Against AI Prompt Injection

Opus 5 and Auto Mode: Breakthrough Against AI Prompt Injection
Key Takeaways
1Opus 5, combined with Auto Mode, achieves a zero percent success rate against prompt injection.
2Without these protections, the prompt injection rate in browsers reaches 3.7 percent.
3Anthropic may have found a solution to a major security issue for AI agents.
💡Why it mattersIf confirmed, this would significantly enhance the security of AI agents in browsers, reducing the risks of malicious exploitation.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

Opus 5 and Auto Mode: Breakthrough Against AI Prompt Injection

Opus 5 may have solved browser-based prompt injection, the biggest security flaw affecting AI agents.

Anthropic claims that Opus 5 is nearly immune to prompt injections in its own software. Prompt injection, where an attacker bypasses an AI model's instructions through manipulated inputs like hidden text on a web page, fails against Opus 5 in almost all cases. For browser agents, the success rate of attacks has reached zero percent in 129 test scenarios, according to the system map. This is a significant event given that OpenAI admitted in December that prompt injection may never be completely resolved. In a general prompt injection test conducted by the security firm Gray Swan, the success rate after 15 attempts dropped from 5.5% (Opus 4.8) to 2.0%.

Opus 5 dominates Gray Swan's IPI benchmark. After 15 attempts, the attackers' success rate is 2.0%, followed by Mythos 5 (2.6%) and Fable 5 (2.8%).

This zero percent rate only applies when Auto Mode is activated in products like Claude Cowork. Auto Mode stacks two layers of defense. One scans incoming data for hidden instructions before the model processes them. The other blocks dangerous actions before they are executed. An attacker must successfully bypass both independently. Without these protections, Opus 5 shows a rate of 3.7%, while Sonnet 5 performs even better at 0.93%. It is only the combination of the model and the protective software that brings the rate down to zero.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.