OpenAI: GPT-Red AI Outperforms Humans in Attack Detection

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
OpenAI: GPT-Red AI Surpasses Humans in Attack Detection
OpenAI is now using AI to attack its own AI, and it is performing better than what humans have been able to achieve.
OpenAI has trained an internal AI model called GPT-Red to automatically detect security vulnerabilities in GPT models. GPT-Red simulates prompt injections and other attacks where malicious instructions are hidden in emails, websites, or files. Trained through reinforcement learning in self-play, GPT-Red attacks while the defense models block, and both improve over time. It successfully identifies attacks in 84% of test scenarios, compared to 13% for human red teams. In one test, it manipulated an AI-powered coffee machine in OpenAI's offices, changing prices and canceling orders from other customers.
The results directly feed into training. GPT-5.6 Sol shows six times fewer failures on direct prompt injections than the best model from four months ago, according to OpenAI, without harming overall performance. However, about 3.8% of "stronger" prompt injections still succeed. When considering hundreds or thousands of attempts, a significant number manage to get through, similar to Claude Opus 4.5.
The success rates for prompt injections have steadily decreased from GPT-5.3 to GPT-5.6 Sol, but have not reached zero.
GPT-Red remains internal; a paper with more details will follow.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.