Anthropic Forced to Disable Fable 5 and Mythos 5 by Washington

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Washington Imposes a Global Suspension on Anthropic
The U.S. government has ordered Anthropic to immediately disable its Fable 5 and Mythos 5 artificial intelligence models for all users worldwide. This directive, motivated by national security concerns, affects not only international clients but also foreign employees of Anthropic, who are also deprived of access to these models. The export ban thus applies to all foreign nationals, whether inside or outside the United States.
While complying with this directive, Anthropic has publicly expressed its reservations, calling the decision a "misunderstanding." The company is actively working to restore access to its models as quickly as possible. Despite this suspension, all other Anthropic models remain available to users.
The Risks of Jailbreaks in Question
The U.S. government justified its decision by discovering a potential "jailbreak" in the Fable 5 and Mythos 5 models, which could allow users to bypass built-in security measures. Anthropic analyzed this allegation and concluded that the identified vulnerabilities are minor and already known, similar to those in other models like OpenAI's GPT-5.5. The company examined a demonstration of this technique and found that it only identifies a "small number of minor vulnerabilities already known" that other publicly available models could also detect.
The potential for jailbreaks involves asking the model to read specific code and correct software bugs. Anthropic asserts that the capabilities demonstrated by the government are "largely available from other models," including OpenAI's GPT-5.5. Security researchers are already using these capabilities daily to protect systems.
A Security Strategy Put to the Test
Before the launch of Fable 5 and Mythos 5, these models underwent rigorous testing by the U.S. government, the UK AI Safety Institute, private third-party organizations, and internal teams. In total, thousands of hours of testing were conducted to assess their safety. Anthropic has highlighted security measures that it believes are more effective than those of previous models. The security measures are "substantially more effective than those of any previously deployed model," states Anthropic. Users have even complained that they are too restrictive.
However, despite these precautions, no model is immune to "jailbreaks." Anthropic has adopted a "defense in depth" strategy to limit these attacks, while acknowledging that perfect resistance is impossible. No tester has found a universal jailbreak, a method that could broadly bypass the model's security measures and unlock a wide range of cyber capabilities. But Anthropic also claims that perfect resistance to jailbreaks is not possible for any model provider currently, a well-documented fact given the high number of attack vectors that LLMs offer.
Knowing this, the company has pursued a strategy it calls "defense in depth": keeping jailbreaks either tightly targeted or costly to execute, combined with broad monitoring to quickly detect and stop successful attacks. Part of this strategy includes a 30-day data retention policy for customer data, which Anthropic says creates "real costs for us with customers" but allows for research and mitigation of jailbreaks.
A Troubling Precedent for the Industry
Anthropic is complying with the order but clearly voices its objections. "We do not agree that the discovery of a narrow potential jailbreak should justify the recall of a commercial model deployed to hundreds of millions of people." If this standard were applied across the industry, it would effectively halt all new deployments of models from every leading model provider, the company argues.
In previous public statements, Anthropic has contended that the government should have the power to block dangerous deployments, but through a legal process that is "transparent, fair, clear, and based on technical facts." The current action does not meet these principles, according to the company, suggesting that this could become another chapter in the ongoing conflict between Anthropic and the U.S. government.
The U.S. government recently issued a new executive order allowing AI developers to submit their models for government security evaluation before release. Anthropic welcomed this approach, but the process was apparently not yet in place when the directive was issued.
The Persistent Challenges of LLMs
Jailbreaks and the related issue of prompt injections remain unresolved security problems since the inception of large language models. No LLM manufacturer is immune. The vulnerability has been known since at least GPT-3 and affects all systems based on LLMs. ChatGPT and Claude can still be attacked by prompt injections under certain conditions, even though their creators have added countermeasures.
Even targeted security efforts have failed. About a year ago, Anthropic built a specialized defense against manipulation attempts and subjected it to a public jailbreaking challenge. After five days, over 300,000 messages, and approximately 3,700 hours of collective work, the system was completely cracked, including a universal jailbreak.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.