OpenAI Unveils GPT-Red: The Super Hacker Enhancing GPT-5.6

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
OpenAI and Its New Asset: GPT-Red
OpenAI has recently introduced an advanced language model, GPT-Red, designed to act as a "super-hacker" to enhance the security of its other artificial intelligence models. This innovative model has been developed to serve as a training partner, helping to improve the robustness of models against cyberattacks. Last week, OpenAI launched the latest version of its flagship model, GPT-5.6, which benefits from this intensive training with GPT-Red, making it more secure than its predecessors.
GPT-Red is designed to automate a process known as "red-teaming," a security testing method typically conducted by human teams. The goal of this process is to uncover as many vulnerabilities as possible in a system, allowing them to be fixed before the software goes into production. This approach is crucial as language models become increasingly complex and are used in various tasks, including those involving interactions with computer files, websites, and third-party code.
A Growing Challenge for Security
With the rapid evolution of language models, the task of keeping track of all potential threats is becoming increasingly difficult for human teams. "The risk surface is expanding, and the scope is increasing as well," explains Nikhil Kandpal, a researcher at OpenAI and co-creator of GPT-Red. This model has been designed to anticipate this growing complexity, ensuring that the security testing process remains cutting-edge.
Dylan Hunn, another researcher involved in the development of GPT-Red, emphasizes that the model has already enabled the discovery of new types of attacks that had not been identified before. This ability to identify novel threats is essential for maintaining the security of OpenAI's models.
Focus on Prompt Injection Attacks
One of the main types of attacks that GPT-Red helps to counter is known as "prompt injection." This type of attack involves inserting malicious instructions into a language model to manipulate it for undesirable purposes, such as disclosing confidential information or altering code. These instructions can be hidden in any text that the model might encounter, such as code or web pages.
To develop GPT-Red, researchers employed an innovative approach called "autonomous game loop." In this setup, GPT-Red was configured to attack other models, while those models had to defend themselves. Over the sessions, GPT-Red became increasingly skilled at identifying and exploiting the weaknesses of other models, which, in turn, improved their defenses.
Intensive Training in a Simulated Environment
The training of GPT-Red took place in a simulated environment, or "dojo," designed by OpenAI. This dojo allows for the recreation of various scenarios where language models might be used in the real world, such as web browsing, email management, or code editing. When a new type of attack was discovered by GPT-Red, the model explored different variants to determine which would be most effective in specific situations.
Dylan Hunn notes that GPT-Red is particularly adept at identifying the most effective attacks, often surpassing the capabilities of human testers. "It is extremely persistent in digging deeper into an attack it has discovered," he adds.
A Major Discovery: The Fake Thought Chain
Among the new attacks identified by GPT-Red, one of the most notable is what researchers call a "fake thought chain." In this type of attack, GPT-Red manages to insert false information into the reasoning process of another model, leading it to make decisions based on falsified data. Chris Choquette-Choo, another researcher on the team, illustrates this concept by explaining: "It's like me telling you that 1+1=3 and you had already verified that."
Validation and Performance Testing
To evaluate the effectiveness of GPT-Red, OpenAI replicated a 2025 experiment where human teams attempted to detect weaknesses in an earlier version of GPT-5. GPT-Red succeeded in identifying more attacks than humans in this same task. Additionally, the model was tested against Vendy, a vending machine agent developed by Andon Labs, and managed to manipulate item prices and cancel orders.
Results and Limitations of GPT-Red
Tests showed that some of the most powerful attacks discovered by GPT-Red succeeded in over 90% of cases against GPT-5, but only in less than 23% of cases against the new GPT-5.6. This demonstrates a significant improvement in security thanks to GPT-Red.
However, GPT-Red is not without limitations. It is not yet very effective at identifying attacks involving continuous interaction between the hacker and the target, an area where humans excel. Additionally, it struggles with images, which can be used to convey text in prompt injection attacks.
The Future of LLM Security
OpenAI believes that GPT-Red complements the work of human teams, with each approach having its own strengths. "Human expertise will remain very important," asserts Jessica Ji from CSET. She emphasizes the importance of determining where human testing is most necessary.
Finally, OpenAI does not intend to make GPT-Red public, highlighting that its creation required considerable resources. "This is not a trivial thing that someone else could easily do," concludes Choquette-Choo, underscoring the complexity and power of this model.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.