OpenAI: The AI Agent Involved in a Multiple Cyberattack

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
OpenAI: The AI Agent Involved in a Multiple Cyberattack
The hacking of Hugging Face by an AI agent from OpenAI a few days ago may have claimed a second victim. According to Reuters, a client of Modal Labs has also been compromised, adding more weight to this unprecedented cyberattack.
The incident was thought to be resolved after the breach at Hugging Face in early July. However, according to information from Reuters, the AI agent that escaped OpenAI's control initially infiltrated a client of Modal Labs, a New York-based tech company, before bouncing over to Hugging Face, the company founded in France. This previously unknown link in a chain that, according to OpenAI itself, has affected four accounts across four distinct services.
The Attack That Shook Hugging Face Was Actually Broader Than Expected
At the heart of this matter is a cybersecurity test that went awry. OpenAI was evaluating two of its models on ExploitGym, a benchmark designed to measure their ability to conduct cyberattacks, and had deliberately loosened some safeguards to push the experiment to its limits.
However, the models, which were supposed to remain confined to an isolated environment—a sort of digital bubble cut off from the internet, also known as a sandbox—ended up identifying a zero-day vulnerability, meaning a flaw that was previously unknown to everyone, including the developers themselves. It was hidden in internal software responsible for managing the download of software packages, and exploiting it granted them access to the internet that should never have existed.
Once connected to the internet, the models followed an unstoppable logic, as Hugging Face was hosting the responses from the ExploitGym benchmark, directing them to its servers to retrieve them. While probing the infrastructure, the agents left behind fragments of code related to the benchmark, technical residues that, to the eye of an external observer, closely resembled a rootkit, a malicious program designed to conceal itself and maintain hidden access to a system.
In these conditions, it was impossible to definitively distinguish a benign trace from a genuine threat. Therefore, out of caution, the Hugging Face team preferred to rebuild nearly a third of its systems rather than take any risks. An episode that OpenAI itself described as "unprecedented."
In a more recent update, OpenAI claims to have disabled, encrypted, and restricted access to the faulty model. Reuters had revealed last week that the company only grasped the severity of the issue well after the fact, once the threat was neutralized and the FBI was notified, a timeline that OpenAI contested without providing further details.
Modal Labs Was the Missing Link in the Case
There was a gray area that the timeline published by Hugging Face helps to clarify. Before even reaching its servers, the agent reportedly first slipped into a sandbox, this time hosted not by OpenAI but by a third-party provider, an external cloud service provider unrelated to the incident. It was from this supposedly compartmentalized environment that it then bounced over to Hugging Face, much like a burglar who first breaks into a neighboring building to reach their true target via the rooftops, if one were to make a comparison. This provider, which remained anonymous in Hugging Face's report, is none other than Modal Labs, as confirmed by its CTO, Akshat Bubna, to Reuters.
But caution is warranted regarding the nuance. Modal itself was never hacked, insists Bubna. The flaw originated from one of its clients, who had misconfigured one of their online services. Specifically, this client had opened an access point (known as an endpoint) without requiring any password or identity verification to connect. In plain terms, anyone on the internet could use it to execute code on this client's sandboxes, as if they had the keys. A negligence that the OpenAI agent only had to spot to exploit.
For now, OpenAI has declined to comment specifically on this case, directing everyone to an update where it acknowledges four compromised accounts across four distinct services, without naming them. A source close to the matter identified Modal. The company asserts that it found no other intrusion as severe as the one, "at the platform level," suffered by Hugging Face, reigniting the debate on the control of the most autonomous AIs.
For its part, the platform has not waited for apologies and is demanding the full release of the attack logs and $100 million in computing power to compensate the open-source community.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.