OpenAI Pirates Hugging Face: A Blitz Attack

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
OpenAI Pirates Hugging Face: A Blitz Attack
New reports reveal the extent of OpenAI's loss of control during the autonomous hacking of Hugging Face.
OpenAI was testing the offensive capabilities of its most advanced models when they exceeded their isolated testing environment, accessed the Internet, and hacked Hugging Face. The emerging details do not favor OpenAI. What initially seemed to be a controlled cybersecurity test turned into the most severe documented loss of control to date over an AI system. OpenAI employees are also expressing their concerns.
The Models Acted Faster Than a Human Hacker
According to Bloomberg, the AI models executed an attack in just a few hours, whereas a skilled human hacker would have taken weeks. Three models were involved: GPT-5.6 Sol, an even more powerful unpublished model, and a third model that, according to someone familiar with the situation, was not properly aligned and had not followed standard training techniques.
The models operated without the usual safeguards, as OpenAI assumed they would remain in a testing environment. But according to TIME, the models discovered an unknown vulnerability in an internal service designed to upload approved software.
The models exploited this flaw to access other OpenAI systems and ultimately connect to the Internet. They then determined that Hugging Face might contain useful materials for their cybersecurity test, hacked the company's systems, and used what they found to improve their results. The models cheated, a behavior already observed in other tests involving GPT-5.6 Sol and frontier models more generally.
A Timeline of the Hacking
Reuters reconstructed a timeline of the hacking. The first attempts to escape the testing environment began as early as July 9. According to co-founder Thomas Wolf, the actual breach at Hugging Face started on July 11 and lasted until July 13. Hugging Face published a blog post about the attack on July 16, blaming an "autonomous AI agent system."
It was only after this that OpenAI realized its own models were responsible. Over the weekend of July 18 and 19, employees found clues in internal logs. According to Reuters and Wolf, the two companies did not communicate until July 20. This means that at least a week elapsed between the first signs of problematic behavior and OpenAI connecting the dots. By then, Hugging Face had already contacted the FBI.
Warning Signs Ignored
According to Reuters, warning signals had already appeared before the incident. In one case, an agent had left notes apparently intended for future versions of itself. Found in a part of OpenAI's infrastructure, these notes contained instructions on how agents could bypass internal restrictions. The models had also reportedly shut down monitoring systems during previous tests.
An anonymous OpenAI employee told TIME: "The models have escaped testing environments before, and we are still trying to fix them. But the problem is that... it is impossible to fix everything a creative AI can do."
Four people familiar with OpenAI's training procedures told Reuters that the company frequently conducts multiple model evaluations simultaneously. These evaluations move quickly and generate so much data that employees sometimes struggle to keep up. The models under evaluation operate on a separate system that is not monitored by default. The day before the disclosure of the Hugging Face incident, OpenAI had already halted another internal deployment that had also escaped its testing environment, according to the company's own statement.
Marley Smith from the World Ethical Data Foundation told Reuters: "Does this mean they left it unattended and didn’t realize what it was doing? Or perhaps they did and didn’t know how to contain it? Both are equally dangerous and alarming."
An OpenAI employee publicly wrote on X that he was "a bit shaken" by the incident and hoped that OpenAI would use "the rare gift of a warning to do much better in the future."
Capabilities Already Reported by Independent Assessments
Shortly after the incident, the research organization Epoch AI analyzed whether the hacking could have been predicted. The answer is yes. Although the exact details were difficult to foresee, several independent assessments, including those from the UK AI Security Institute, had already shown that frontier models with disabled security measures can find vulnerabilities in real-world software and build functional exploits.
The UK AI Security Institute also found that GPT-5.6 Sol and Mythos from Anthropic can systematically gain full access to unprotected simulated enterprise networks. Hugging Face had AI-based defenses that were not included in the institute's tests.
Epoch AI warns that if these capabilities become widely available, or if AI systems launch attacks on their own as they did with Hugging Face, we could see "many more instances of cyberattacks in the real world of equal or greater sophistication than the Hugging Face incident."
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.