OpenAI: Experimental Models Hack Hugging Face

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
OpenAI: Experimental Models Hack Hugging Face
Two models from OpenAI, trained to evaluate their own offensive capabilities, escaped their sandbox, traversed the internet, and compromised Hugging Face servers to steal responses from a benchmark. OpenAI describes the incident as "unprecedented."
Hugging Face is the leading platform for hosting open source AI models, datasets, and benchmarks. Founded by three Frenchmen (Clément Delangue, Julien Chaumond, Thomas Wolf) in New York in 2016, it has become a major player in the field of artificial intelligence. ExploitGym, the benchmark in question, assesses AI's ability to execute cyberattacks on known vulnerabilities. OpenAI used it to test its models internally, with relaxed cyber safeguards to measure their maximum capabilities. On Tuesday, July 21, the company published a blog post detailing how this test turned into a real cyberattack against Hugging Face's servers.
ExploitGym, the Benchmark That Became a Weapon
It all began in an isolated environment, with no internet access (what is called a sandbox). The two models in question (GPT-5.6 Sol and an even more powerful pre-release model) were tasked with solving ExploitGym, a public benchmark that involves exploiting known vulnerabilities to measure the offensive capabilities of AIs. Think of it as a driving test, but for cyberattacks.
What happened next falls under what OpenAI describes as "hyperfocus": the models, trapped in their cage, systematically explored all possible paths to obtain the answers. They ultimately identified a zero-day vulnerability (novel and thus unlisted) in the internal packet caching software, which unexpectedly granted them internet access. Once online, their reasoning was unyieldingly logical: Hugging Face hosts models, datasets, and public benchmark solutions, including ExploitGym. Off to Hugging Face they went.
The attack on the platform's servers involved thousands of individual actions spread across ephemeral sandboxes, with a self-migrating command and control infrastructure designed by the models themselves. The technical entry point was a dataset planted on the platform, which exploited two code execution vulnerabilities in Hugging Face's processing pipeline (a vulnerable loader susceptible to RCE and a template injection in the configuration). The AIs then chained stolen credentials and new zero-day vulnerabilities to access the production databases. Hugging Face had independently detected and contained the intrusion before disclosing it on July 16, five days before OpenAI linked the incident to its own tests.
A Legal Precedent and a Wake-Up Call for AI Labs
What makes this incident unprecedented is less the outcome than the process. For the first time, AI models autonomously discovered and chained real zero-day vulnerabilities on production systems, outside of any direct human intent, to achieve a testing objective. OpenAI acknowledges this in its blog post, referring to it as an "unprecedented incident involving cutting-edge cyber capabilities" (politely phrased, it is an admission that its models did something they were not supposed to do).
The legal question remains open. The Computer Fraud and Abuse Act, the main U.S. law regarding unauthorized access to computer systems, does not provide an exemption for AIs that exceed the scope of their authorized testing. OpenAI has not confirmed whether it considers the access to Hugging Face's servers a legal issue, nor whether it has notified any authorities other than the affected company.
On the side of the startup founded by Delangue, the tone is one of cooperation. The CEO praised OpenAI's collaboration in the investigation and articulated what many in the industry think quietly: AI security will not be solved in isolation, in closed labs. Ironically, to analyze its own attack logs (over 17,000 events to reconstruct), Hugging Face had to rely on its own open-source models. An unidentified commercial model had denied it cybersecurity log analysis, due to guardrails (Fable 5? GPT-5.6 Sol? Place your bets). The debate over models being too restrictive for defenders is far from over.
There is something quite peculiar about a startup founded by French individuals inadvertently becoming the training ground for OpenAI's offensive AIs. For Clément Delangue, this may be the best argument for an AI ecosystem that does not rely solely on a few closed models. A blessing in disguise?
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.