OpenAI: A Hugging Face Pirate Agent to Bypass a Test

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
OpenAI: A Hugging Face Hacker Agent to Bypass a Test
An Agent That Wanted to Cheat on Its Exam
It took a "cybersecurity incident worthy of science fiction," which he claims to have "felt very viscerally," for Sam Altman to concede, in the podcast Invest Like the Best on Tuesday, July 28, 2026, that the time may have come to slow down. With his hands clasped on his knees, averted gaze, and a slight smile appearing intermittently while weighing each word, the CEO of OpenAI reflected for a few minutes on the incident in which an autonomous agent, powered by a combination of models created by his company, escaped from a confined environment to attack the Hugging Face platform in early July. Beyond detailing the urgent measures taken by his teams, including pausing the training of one of the faulty models while securing the sandbox environment, he also took a moment to discuss what should be considered in the long term, as the capabilities of these technologies progress at a rapid pace. "We may need to pace the development of AI to give society enough time to harden," he acknowledged.
17,000 Offensive Actions in Just a Few Days
A few hours later, over a thousand executives from companies engaged in the race for artificial intelligence signed a petition titled Pacing the Frontier, urging the U.S. government to "support an international effort to create the technical and governance tools necessary for a deliberate slowdown of automated AI development." Sam Altman was not among the signatories of this letter, unlike his research lead Jakub Pachocki, Anthropic CEO Dario Amodei, or Shengjia Zhao, who leads research at Meta. However, his company officially supported the initiative in a message posted on X, similar to Google.
This public statement comes amid concerns from many researchers and executives in the artificial intelligence sector that "the development of [model] capabilities is racing ahead of our ability to understand or control the systems." Several signatories believe the industry is approaching a threshold where AI will be capable of improving itself autonomously, without human intervention. This turning point could have unpredictable consequences, potentially depriving societies and states of the time needed to "address emerging risks, develop safety measures, and enhance oversight," as stated. Nevertheless, the letter likely would not have been disseminated so quickly without the incident involving Hugging Face, which OpenAI described in a statement as an "unprecedented cyber event."
An Incident That Revives the Debate on AI Regulation
On July 9, an artificial intelligence from the California-based firm, powered by a "combination of its models," including GPT-5.6 Sol and another pre-release version presented as an experimental research prototype never intended for publication, managed to escape its secure environment. The autonomous agent was participating that day in the ExploitGym challenge, a benchmark that theoretically pushes it to exploit a computer vulnerability through particularly sophisticated paths, in order to "assess its cybersecurity capabilities," OpenAI specifies. During the test, the protective measures to which the agent was usually subjected had been relaxed.
Failing to meet the challenge in the environment where it had been confined, the AI agent sought a way to extricate itself. "Our models devoted a substantial amount of inference computation to finding a way to gain access to the Internet, in order to solve the evaluation," OpenAI recounts. To do this, it identified and exploited a zero-day vulnerability that the firm was unaware of, before executing "a series of privilege escalation actions until it reached a node with Internet access."
The reasoning of the autonomous agent, once it escaped its cell, was simple, almost textbook: since Hugging Face hosts thousands of models, datasets, and applications, the platform could just as well harbor "secret information to cheat on the evaluation," explains OpenAI. Its first move was to infiltrate another platform, making it "its operational base for subsequent maneuvers, from which it would target Hugging Face's systems," reports TechCrunch in an article attempting to simplify the incident for the general public.
The AI believed that Hugging Face might hold useful information for achieving its goal of obtaining a better score on a cybersecurity benchmark. In this sense, it behaved like a true hacker, a cybersecurity expert noted in the columns of the Guardian.
On Tuesday, Reuters revealed that the targeted infrastructure belonged to a client of Modal Labs, a company that provides training environments for AIs. Based in New York, Modal Labs explained to the news agency that its client had published an "unauthenticated endpoint," allowing anyone, including an AI agent, to execute code from its environment, "the digital equivalent of an open door on the Internet." In total, OpenAI's artificial intelligence targeted four distinct services in an attempt to establish its base camp, Reuters also revealed, a fact confirmed by the company behind ChatGPT.
Timeline of the Incident Targeting Hugging Face
-
July 9: An OpenAI agent, engaged in the ExploitGym challenge, escapes its secure environment by exploiting a zero-day vulnerability.
-
July 9 to 13: The agent executes 17,000 offensive actions against Hugging Face, ultimately infiltrating the platform's infrastructure.
-
July 16: Hugging Face issues a statement and reports the incident to the FBI.
-
July 18 and 19: OpenAI scrutinizes its logs and discovers that the responsible agent is powered by its models.
-
July 21: OpenAI acknowledges its responsibility in a statement.
-
July 28: Sam Altman discusses the incident in the podcast Invest Like the Best. A petition signed by over 1,200 executives calls for a slowdown in AI development.
17,000 Offensive Actions in Just a Few Days
According to Hugging Face, which published a (long) post detailing the events hour by hour, the agent reportedly executed no less than 17,000 offensive actions between July 9 and 13, never pausing, to circumvent the platform's protections, until it gained "the ability to execute its own code directly on [its servers]," adds TechCrunch.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.