Brief IA

Claude Fable 5: Anthropic Confronts Hidden Safeguards

🤖 Models & LLM·Tom Levy·

Claude Fable 5: Anthropic Confronts Hidden Safeguards

Claude Fable 5: Anthropic Confronts Hidden Safeguards
Key Takeaways
1Claude Fable 5, equipped with Mythos capabilities, has sparked debates about the transparency of its hidden safeguards.
2Cybersecurity experts are concerned about restrictions that could hinder defenders while blocking attackers.
3Anthropic has responded by promising to make functionality degradations visible to users.
💡Why it mattersManaging security in advanced AIs raises crucial questions about the balance between protection and innovation.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

Claude Fable 5 and Its Hidden Safeguards: A Question of Transparency

The introduction of Claude Fable 5 by Anthropic has sparked a wave of reactions online, not due to its technological power, but rather because of the hidden safeguards that accompany it. While this tool offers Mythos-class capabilities, users quickly discovered that these secret security measures raised trust issues.

Reactions to the Controversy

The controversy surrounding Fable 5 does not focus on the power of the AI itself, but on the transparency of the integrated security measures. The safeguards, which were initially intended to protect users, have led researchers to doubt the exact nature of the tests they were conducting. Cybersecurity experts have expressed concern about the potential impact of these guardrails, which could not only block attackers but also hinder defenders' efforts.

Mythos, introduced in April as part of the Project Glasswing, is the result of a collaboration between major tech organizations and Anthropic. This project aims to identify and rectify vulnerabilities in Internet infrastructure. However, due to its ability to detect unknown flaws, access to this tool has been restricted to certain organizations, as it could also be used to exploit these vulnerabilities.

Communication Issues

Anthropic had clearly stated that Fable would not support certain research deemed risky in fields like cybersecurity, biology, and chemistry. Despite this, some experts have warned against blind trust in the company's security claims. Sally Vincent, a threat research engineer at Exabeam, emphasized in an email that claims of resistance to jailbreaks should be taken with caution, as the results only represent a snapshot in time. She added that attackers are constantly adapting.

The Silent Degradation

For researchers involved in ambitious projects, such as designing ultra-powerful chips or developing cutting-edge AI language models, Fable operated quietly. Like other reported projects, it degraded Fable models to Opus without informing users. A mention in the 319-page system card for Fable and Mythos indicated that this degradation would occur when working on such types of projects, but the behavior remained invisible to those who had not taken the time to read the document in full.

Expert Reactions

Rob T. Lee, head of AI at the SANS Institute, expressed concerns about the restrictions imposed by Fable, which prevent defenders from developing new defensive capabilities. He noted that while these measures are intelligent, they hinder progress in creating innovative defenses.

Anthropic's Response

In response to the online reactions, Anthropic quickly announced changes to the safeguards of Fable 5. The company is committed to making functionality degradations visible to users. Starting this week, reported requests will visibly be downgraded to Opus 4.8. Additionally, on the API, any reported request will provide a reason for its denial.

Anthropic clarified that its current set of safeguards covers a limited number of tasks, such as cutting-edge LLM data pipelines and the development of kernels for certain non-standard chips.

Concerns About False Positives

Anthropic also acknowledged concerns related to false positives. Currently, the classifier triggers on about 0.05% of tasks, affecting less than 0.05% of organizations. To be more robust, a visible safeguard must be broader, which leads to incorrectly flagged requests.

Conclusion

Anthropic admitted to having made a mistake in balancing its security measures and is committed to reducing false positives as quickly as possible. The company also explained that the hidden safeguards were designed to be more difficult to circumvent, although their presence was quickly discovered.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.