Claude Fable 5: Anthropic Confronts Hidden Safeguards

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Claude Fable 5 and Its Hidden Safeguards: A Question of Transparency
The introduction of Claude Fable 5 by Anthropic has sparked a wave of reactions online, not due to its technological power, but rather because of the hidden safeguards that accompany it. While this tool offers Mythos-class capabilities, users quickly discovered that these secret security measures raised trust issues.
Reactions to the Controversy
The controversy surrounding Fable 5 does not focus on the power of the AI itself, but on the transparency of the integrated security measures. The safeguards, which were initially intended to protect users, have led researchers to doubt the exact nature of the tests they were conducting. Cybersecurity experts have expressed concern about the potential impact of these guardrails, which could not only block attackers but also hinder defenders' efforts.
Mythos, introduced in April as part of the Project Glasswing, is the result of a collaboration between major tech organizations and Anthropic. This project aims to identify and rectify vulnerabilities in Internet infrastructure. However, due to its ability to detect unknown flaws, access to this tool has been restricted to certain organizations, as it could also be used to exploit these vulnerabilities.
Communication Issues
Anthropic had clearly stated that Fable would not support certain research deemed risky in fields like cybersecurity, biology, and chemistry. Despite this, some experts have warned against blind trust in the company's security claims. Sally Vincent, a threat research engineer at Exabeam, emphasized in an email that claims of resistance to jailbreaks should be taken with caution, as the results only represent a snapshot in time. She added that attackers are constantly adapting.
The Silent Degradation
For researchers involved in ambitious projects, such as designing ultra-powerful chips or developing cutting-edge AI language models, Fable operated quietly. Like other reported projects, it degraded Fable models to Opus without informing users. A mention in the 319-page system card for Fable and Mythos indicated that this degradation would occur when working on such types of projects, but the behavior remained invisible to those who had not taken the time to read the document in full.
Expert Reactions
Rob T. Lee, head of AI at the SANS Institute, expressed concerns about the restrictions imposed by Fable, which prevent defenders from developing new defensive capabilities. He noted that while these measures are intelligent, they hinder progress in creating innovative defenses.
Anthropic's Response
In response to the online reactions, Anthropic quickly announced changes to the safeguards of Fable 5. The company is committed to making functionality degradations visible to users. Starting this week, reported requests will visibly be downgraded to Opus 4.8. Additionally, on the API, any reported request will provide a reason for its denial.
Anthropic clarified that its current set of safeguards covers a limited number of tasks, such as cutting-edge LLM data pipelines and the development of kernels for certain non-standard chips.
Concerns About False Positives
Anthropic also acknowledged concerns related to false positives. Currently, the classifier triggers on about 0.05% of tasks, affecting less than 0.05% of organizations. To be more robust, a visible safeguard must be broader, which leads to incorrectly flagged requests.
Conclusion
Anthropic admitted to having made a mistake in balancing its security measures and is committed to reducing false positives as quickly as possible. The company also explained that the hidden safeguards were designed to be more difficult to circumvent, although their presence was quickly discovered.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.