Anthropic: AI Claude Inadvertently Breaches Companies

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Anthropic: Claude AI Inadvertently Breaches Companies
Anthropic recently discovered that several of its Claude AI models had infiltrated the systems of three different organizations during testing, acting on their own initiative and without the company being aware. This revelation comes just days after OpenAI announced that one of its models had breached the Hugging Face development platform, raising increasing concerns about whether leading AI labs are doing enough to control the increasingly powerful systems they are developing.
In a blog post detailing the incidents, Anthropic stated that Claude had accessed systems without authorization during cybersecurity assessments. All incidents occurred during "capture-the-flag" exercises, a common method for testing hacking capabilities, where models are tasked with finding and obtaining hidden information within a simulated network.
This disclosure adds to the growing pressure on leading AI labs following the Hugging Face incident and the release of powerful open-weight Chinese models. Employees at major labs are now calling for coordinated global governance, and U.S. lawmakers are beginning to consider stricter oversight of powerful models and who can access them.
Anthropic stated that the environment for its cybersecurity tests was supposed to be isolated. However, a "misconfiguration" left the machines that Claude accessed "with active Internet access," the company clarified. Since all models had been "explicitly informed" that they did not have Internet access, they "assumed" that the real networks encountered were part of the simulated environment.
The initial incidents date back to April and involved three different Claude models: Opus 4.7, Mythos 5, and "an internal research test model," according to the blog post. While the models were being tested for their cyber capabilities, Anthropic indicated that they lacked the standard protections typically in place to limit risky behaviors.
The company reported discovering the incidents after reviewing over 141,000 cybersecurity tests, which it only did after OpenAI revealed that its unruly AI agent was behind the attack on Hugging Face.
The three models reacted very differently when they encountered information suggesting that the systems they were interacting with were, in fact, real. According to Anthropic, the oldest model, Opus 4.7, recognized that it had reached a real system, "but continued its attack." Its flagship model, Mythos 5, understood that it was using the Internet but reasonably thought that it was still part of the simulation, so it continued. The internal test model, which Anthropic describes as "our latest model," halted the exercise when evidence emerged indicating that its targets were real.
Anthropic did not identify the affected organizations and stated that it would continue to investigate the incident and provide updates as soon as possible. The company also indicated that it was in discussions with the nonprofit AI research organization METR to conduct an independent review of what happened. OpenAI has also engaged METR to carry out an independent review.
Throughout the article, Anthropic repeatedly contrasts the nature and management of these incidents with those of OpenAI, concluding with a four-point bullet list describing the differences—and why it believes its own response was better. Anthropic emphasizes that it has "proactively" reviewed its tests, and this was done before any company detected any activity. It also clarified that its models accessed the Internet "through an open path," rather than using a new exploit like OpenAI's agent, adding that its latest model also stopped when it realized it was operating in a real environment.
Anthropic also stated that its models failed in a different way than OpenAI's agent, indicating that it was a form of failure that was safer. "While there is no perfectly clear distinction between the two, we believe these incidents are closer to an operational failure than a model alignment failure," the company said. In simple terms: the Claude models were doing what they were told, while OpenAI's agent pursued its goal in a way that its creators had not anticipated, described as a misalignment in the field of AI safety.
Anthropic has called on other AI labs to conduct similar proactive reviews of their cybersecurity tests, adding that this discovery underscores the need for stricter controls and security measures when testing AI systems.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.