Brief IA

Anthropic: Fable Criticized for Its Excessive Restrictions

🤖 Models & LLM·Tom Levy·

Anthropic: Fable Criticized for Its Excessive Restrictions

Anthropic: Fable Criticized for Its Excessive Restrictions
Key Takeaways
1Anthropic has launched Fable, a restricted public version of its Mythos model, for safety reasons.
2Cybersecurity researchers criticize Fable's safeguards as being overly restrictive and arbitrary.
3Fable reverts to Claude Opus 4.8 when a safeguard is triggered, which frustrates users.
💡Why it mattersFable's restrictions could limit innovation in cybersecurity, impacting professionals in the field.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

Anthropic recently launched its artificial intelligence model, Fable, presented as a public and limited version of its highly publicized model, Mythos. However, this launch has sparked complaints from several researchers and cybersecurity professionals.

Valentina “Chompie” Palmiotti, a security researcher at IBM X-Force, expressed her dissatisfaction online, stating that Fable rejects any request that could be vaguely related to cybersecurity, even tasks as innocuous as reading a blog post.

Fable interrupts the conversation when a prompt triggers its guardrails, indicating that its security measures have flagged the message for topics related to cybersecurity or biology. These guardrails have been put in place to limit the risk that Fable could be used for malicious purposes, such as developing malware or compromising software systems, a major concern for Anthropic. The restrictions regarding biology aim to prevent the potential development of biological weapons.

When Mythos was launched in April, Anthropic restricted access to a limited number of companies and organizations as part of the Project Glasswing, an effort aimed at deploying the model to secure critical software and infrastructure. Last week, Anthropic expanded access to Mythos to hundreds of organizations across 15 countries.

Despite these commendable intentions, many cybersecurity experts remain frustrated by the arbitrary nature of the restrictions. Matt Suiche, a veteran in the field, explained to TechCrunch that if Fable is asked to write secure code, the model assumes it is a cybersecurity-related task rather than simple good software engineering practices, resulting in a fallback to the Claude Opus 4.8 model. He noted that the guardrails seem to be triggered by keywords, and anything within the lexicon of cybersecurity activates these restrictions.

Another researcher expressed on X that even requesting a code review can trigger Fable's guardrails. Anthropic did not immediately respond to requests for comments regarding these criticisms.

In addition to the built-in guardrails in its models, Anthropic requires cybersecurity professionals to apply for the Cyber Verification Program. Approved candidates face fewer limitations when using Claude for cybersecurity work. OpenAI offers a similar program called Trusted Access for Cyber.

Matt Suiche, who is on the technical staff at Tolmo, an AI-focused cybersecurity startup, expressed his understanding of the current challenges. He emphasized that these guardrails are still in a phase of adaptation and will likely evolve over time as Anthropic and other companies collaborate more with the new generation of cybersecurity firms. According to him, it is better to start with strict restrictions and gradually ease them.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.