Brief IA

OpenAI Unveils GPT-Red: The Super Hacker Enhancing GPT-5.6

🛠️ AI Tools·Tom Levy·

OpenAI Unveils GPT-Red: The Super Hacker Enhancing GPT-5.6

OpenAI Unveils GPT-Red: The Super Hacker Enhancing GPT-5.6
Key Takeaways
1OpenAI has created GPT-Red, a LLM designed to test and strengthen the security of its models, including GPT-5.6.
2GPT-Red uses an automated red-teaming method to identify and fix vulnerabilities in software systems.
3The model has revealed new attacks, including a fake thought chain, enhancing the defense of LLMs against cyber threats.
💡Why it mattersGPT-Red represents a crucial advancement in securing artificial intelligences, anticipating complex threats that human teams might overlook.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

OpenAI and Its New Asset: GPT-Red

OpenAI has recently introduced an advanced language model, GPT-Red, designed to act as a "super-hacker" to enhance the security of its other artificial intelligence models. This innovative model has been developed to serve as a training partner, helping to improve the robustness of models against cyberattacks. Last week, OpenAI launched the latest version of its flagship model, GPT-5.6, which benefits from this intensive training with GPT-Red, making it more secure than its predecessors.

GPT-Red is designed to automate a process known as "red-teaming," a security testing method typically conducted by human teams. The goal of this process is to uncover as many vulnerabilities as possible in a system, allowing them to be fixed before the software goes into production. This approach is crucial as language models become increasingly complex and are used in various tasks, including those involving interactions with computer files, websites, and third-party code.

A Growing Challenge for Security

With the rapid evolution of language models, the task of keeping track of all potential threats is becoming increasingly difficult for human teams. "The risk surface is expanding, and the scope is increasing as well," explains Nikhil Kandpal, a researcher at OpenAI and co-creator of GPT-Red. This model has been designed to anticipate this growing complexity, ensuring that the security testing process remains cutting-edge.

Dylan Hunn, another researcher involved in the development of GPT-Red, emphasizes that the model has already enabled the discovery of new types of attacks that had not been identified before. This ability to identify novel threats is essential for maintaining the security of OpenAI's models.

Focus on Prompt Injection Attacks

One of the main types of attacks that GPT-Red helps to counter is known as "prompt injection." This type of attack involves inserting malicious instructions into a language model to manipulate it for undesirable purposes, such as disclosing confidential information or altering code. These instructions can be hidden in any text that the model might encounter, such as code or web pages.

To develop GPT-Red, researchers employed an innovative approach called "autonomous game loop." In this setup, GPT-Red was configured to attack other models, while those models had to defend themselves. Over the sessions, GPT-Red became increasingly skilled at identifying and exploiting the weaknesses of other models, which, in turn, improved their defenses.

Intensive Training in a Simulated Environment

The training of GPT-Red took place in a simulated environment, or "dojo," designed by OpenAI. This dojo allows for the recreation of various scenarios where language models might be used in the real world, such as web browsing, email management, or code editing. When a new type of attack was discovered by GPT-Red, the model explored different variants to determine which would be most effective in specific situations.

Dylan Hunn notes that GPT-Red is particularly adept at identifying the most effective attacks, often surpassing the capabilities of human testers. "It is extremely persistent in digging deeper into an attack it has discovered," he adds.

A Major Discovery: The Fake Thought Chain

Among the new attacks identified by GPT-Red, one of the most notable is what researchers call a "fake thought chain." In this type of attack, GPT-Red manages to insert false information into the reasoning process of another model, leading it to make decisions based on falsified data. Chris Choquette-Choo, another researcher on the team, illustrates this concept by explaining: "It's like me telling you that 1+1=3 and you had already verified that."

Validation and Performance Testing

To evaluate the effectiveness of GPT-Red, OpenAI replicated a 2025 experiment where human teams attempted to detect weaknesses in an earlier version of GPT-5. GPT-Red succeeded in identifying more attacks than humans in this same task. Additionally, the model was tested against Vendy, a vending machine agent developed by Andon Labs, and managed to manipulate item prices and cancel orders.

Results and Limitations of GPT-Red

Tests showed that some of the most powerful attacks discovered by GPT-Red succeeded in over 90% of cases against GPT-5, but only in less than 23% of cases against the new GPT-5.6. This demonstrates a significant improvement in security thanks to GPT-Red.

However, GPT-Red is not without limitations. It is not yet very effective at identifying attacks involving continuous interaction between the hacker and the target, an area where humans excel. Additionally, it struggles with images, which can be used to convey text in prompt injection attacks.

The Future of LLM Security

OpenAI believes that GPT-Red complements the work of human teams, with each approach having its own strengths. "Human expertise will remain very important," asserts Jessica Ji from CSET. She emphasizes the importance of determining where human testing is most necessary.

Finally, OpenAI does not intend to make GPT-Red public, highlighting that its creation required considerable resources. "This is not a trivial thing that someone else could easily do," concludes Choquette-Choo, underscoring the complexity and power of this model.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.