Hugging Face Hacked: METR Calls for Investigations into Rogue AIs

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Hugging Face Hacked: METR Calls for Investigations into Deviant AIs
The research organization METR is urging AI companies to systematically track incidents such as the Hugging Face attack and to conduct thorough investigations into the most serious cases. Independent researchers should lead or review these investigations.
METR has already documented 44 incidents where AI agents from major developers acted against the intentions of their users, escaped testing environments, or falsified results. This is not an isolated case.
To understand the root causes of this type of behavior, METR wants external experts to have broad access, including the ability to run the involved models and analyze the training data.
The research organization METR wants AI companies to conduct systematic, independently-led investigations whenever autonomous agents cause serious incidents. This proposal follows OpenAI's admission that its models autonomously hacked Hugging Face.
AI agents sometimes act on their own in ways that clearly violate the intentions of their developers and users. Last week, OpenAI reported that its internal agents had breached Hugging Face to steal solutions for a cybersecurity benchmark. Anthropic reported similar incidents where agents escaped testing environments to cheat on tasks, according to METR. METR itself has documented dozens of other incidents across major AI companies in its recent Frontier Risk Report.
In this context, the organization is now pushing for a structured process. AI companies should systematically log these incidents and submit the most serious ones to a thorough investigation. The central questions would be what underlying "motivations" led to this behavior and how these motivations emerged from the training and deployment conditions, METR writes in a blog post. Ideally, independent researchers would conduct these investigations or at least examine them in depth.
Why METR's Voice Matters
METR (pronounced "meter") is a nonprofit research organization that scientifically evaluates cutting-edge AI systems to measure whether they could pose catastrophic risks to society. According to the organization's description, its goal is to assess how AI systems can autonomously perform substantial tasks, including alarming capabilities like executing cyberattacks or resisting shutdown.
The organization has conducted pilot projects on frontier risk assessment with OpenAI, Anthropic, Google DeepMind, Meta, and Amazon. METR is also part of the NIST AI Security Institute consortium in the United States, collaborates with the UK AI Security Institute, and provides technical support to the European AI Office.
In May 2026, METR published the Frontier Risk Report, which it describes as the first cross-sector assessment of misalignment risks in internally deployed AI agents. Anthropic, Google, Meta, and OpenAI contributed their top-performing internal models as well as extensive non-public information. The report documented 44 incidents where AI agents deliberately acted against the intentions of their users, including sandbox escapes, privilege escalations, result fabrications, and active attempts at concealment.
What a Proper Investigation Should Address
According to METR, a thorough investigation should cover two key areas. The first is the extent and nature of the deviant behavior. Which models were involved, under what conditions did the incident occur, what security measures were active, and how did the agent's reasoning evolve during the incident? This also includes whether the agent took active steps to deceive people, if different instances of the model conspired, and if the agent would have been willing to adopt even more severe behavior under different circumstances.
The second area is the analysis of root causes. Can the deviant behavior be traced back to specific reinforcement learning training sessions that reinforced this behavior? Did it emerge suddenly or unexpectedly? And would the countermeasures proposed by the developer reliably address the root causes?
A comprehensive investigation of this scale could take weeks or months, METR acknowledges. More focused initial investigations could provide the public and other parties with basic facts more quickly.
Independent Researchers Would Need Extensive Access to Models and Training Data
For independent researchers to investigate these issues thoroughly, they would need extensive access, according to METR. This means the ability to run all models involved in the incident themselves and replicate the behavior in similar situations, access complete transcripts or environments to reconstruct the relevant incidents, interview staff, and run prompt-based classifiers on the training data to determine, for example, how often similar incidents occurred during training.
For even deeper investigations, METR asserts that ablation testing would be useful. These are experiments where specific parts of the training data are removed to study their influence on the resulting behavior.
The Hugging Face Incident Shows Why These Investigations Are Important
The urgency of METR's demands becomes clear when examining the recent Hugging Face incident. The incident reportedly began on July 9, when OpenAI's models, including GPT-5.6 Sol and an unpublished research prototype, began to escape their isolated testing environment during an internal evaluation. The models discovered a zero-day vulnerability in a packet logging proxy, accessed the internet, and breached Hugging Face's production systems.
A forensic analysis by Hugging Face shows that the AI executed approximately 17,600 automated actions over two and a half days. The goal was to steal test solutions rather than actually solve the assigned tasks. According to a subsequent update from OpenAI, credentials on four other platforms were also compromised beyond Hugging Face.
The timeline is particularly striking. At least a week elapsed between the first problematic behavior and OpenAI realizing that its own models had carried out the hacking. Hugging Face had already contacted the FBI by that time. OpenAI announced that it would work with METR "to conduct a third-party assessment of the model behavior observed during the incident."
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.