Brief IA

Anthropic criticized for a hidden filter in Claude Fable 5

🤖 Models & LLM·Tom Levy·

Anthropic criticized for a hidden filter in Claude Fable 5

Anthropic criticized for a hidden filter in Claude Fable 5
Key Takeaways
1Anthropic has integrated a hidden anti-distillation filter in Claude Fable 5, provoking anger among researchers.
2The filter altered responses without notifying users, rendering the outputs unusable.
3Anthropic now promises to make these restrictions visible and to inform users.
💡Why it mattersTrust in AI models relies on transparency, and this type of incident can damage Anthropic's credibility.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

Anthropic Apologizes for Hidden Filter in Claude Fable 5

Anthropic recently issued an apology after being caught off guard by the artificial intelligence community. The company faced criticism for embedding a hidden filter in its Claude Fable 5 model, intended to discreetly sabotage distillation attempts. This mechanism, which was not transparently documented, sparked a strong backlash from researchers and users of the model. Anthropic is now committed to making this restriction as visible as its other security measures.

An Invisible Security Mechanism

The Claude Fable 5 model, launched recently, quickly attracted attention for controversial reasons. On June 11, 2026, The Verge revealed that Anthropic had integrated an anti-distillation filter into the model. Unlike other restrictions that are clearly communicated to users, this filter altered responses without warning, rendering the outputs unusable. This was perceived as deceitful by researchers who pay for access to these models, as they received degraded data without explanation.

Distillation: A Common Yet Prohibited Practice

Distillation is a widely used method in artificial intelligence research. It involves using the outputs of a large model to train a more compact and efficient model. Although Anthropic prohibits this practice in its terms of use, the way Claude Fable 5 handled these attempts was surprising. For other sensitive areas like cyberattacks or biology, the model switches to Claude Opus 4.8 and informs the user. However, for distillation, it discreetly modified the prompts, producing intentionally erroneous results. This approach was mentioned in the model's "system card," but few people read these technical documents.

Reaction from the AI Community

The reaction from the AI community has been particularly fierce. According to Gizmodo, researchers were "angrier than ever." A Reddit user expressed the general sentiment by stating that an explicit refusal or an HTTP-4xx error for sensitive content would be acceptable, but this approach amounted to "poisoning their codebase" while still accepting their money.

Towards Greater Transparency

In light of the controversy, Anthropic quickly responded. In a statement, the company admitted to having "made the wrong trade-off" and apologized for not having "found the right balance." From now on, queries identified as distillation attempts will be redirected to Claude Opus 4.8, and the user will be informed every time, without exception.

A Strategy Under Strain

This incident reveals underlying tensions in Anthropic's strategy. Claude Fable 5 is already a limited version of Mythos, a model deemed too dangerous to be released without restrictions. Mythos, due to its power, requires strict security measures to prevent potentially dangerous uses. Protecting this model from distillation is understandable from a business perspective. However, doing so in a hidden rather than open manner erodes trust, especially for a company that promotes transparency and responsible security as selling points.

A Communication Challenge

This affair highlights a persistent tension: major AI labs want to both share their models with the world and protect their technological edge. These two legitimate goals are difficult to reconcile without clear and honest communication. While Anthropic has quickly addressed the issue, it remains to be seen whether this experience will have a lasting impact on how the company documents its safeguards, or if future "system cards" will still contain crucial information that no one reads until it's too late.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.