Brief IA

OpenAI: GPT-Red AI Outperforms Humans in Attack Detection

🤖 Models & LLM·Tom Levy·

OpenAI: GPT-Red AI Outperforms Humans in Attack Detection

OpenAI: GPT-Red AI Outperforms Humans in Attack Detection
Key Takeaways
1OpenAI has developed GPT-Red, an internal model capable of detecting 84% of attacks during tests, outperforming humans.
2Human testing teams achieved only a 13% success rate, demonstrating the superior effectiveness of AI.
3The results obtained by GPT-Red contribute to the improvement of models like GPT-5.6 Sol.
💡Why it mattersThe use of AI to test and strengthen other AIs could revolutionize the security and reliability of AI systems in the future.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

OpenAI: GPT-Red AI Surpasses Humans in Attack Detection

OpenAI is now using AI to attack its own AI, and it is performing better than what humans have been able to achieve.

OpenAI has trained an internal AI model called GPT-Red to automatically detect security vulnerabilities in GPT models. GPT-Red simulates prompt injections and other attacks where malicious instructions are hidden in emails, websites, or files. Trained through reinforcement learning in self-play, GPT-Red attacks while the defense models block, and both improve over time. It successfully identifies attacks in 84% of test scenarios, compared to 13% for human red teams. In one test, it manipulated an AI-powered coffee machine in OpenAI's offices, changing prices and canceling orders from other customers.

The results directly feed into training. GPT-5.6 Sol shows six times fewer failures on direct prompt injections than the best model from four months ago, according to OpenAI, without harming overall performance. However, about 3.8% of "stronger" prompt injections still succeed. When considering hundreds or thousands of attempts, a significant number manage to get through, similar to Claude Opus 4.5.

The success rates for prompt injections have steadily decreased from GPT-5.3 to GPT-5.6 Sol, but have not reached zero.

GPT-Red remains internal; a paper with more details will follow.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.