Anthropic criticized for a hidden filter in Claude Fable 5

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Anthropic Apologizes for Hidden Filter in Claude Fable 5
Anthropic recently issued an apology after being caught off guard by the artificial intelligence community. The company faced criticism for embedding a hidden filter in its Claude Fable 5 model, intended to discreetly sabotage distillation attempts. This mechanism, which was not transparently documented, sparked a strong backlash from researchers and users of the model. Anthropic is now committed to making this restriction as visible as its other security measures.
An Invisible Security Mechanism
The Claude Fable 5 model, launched recently, quickly attracted attention for controversial reasons. On June 11, 2026, The Verge revealed that Anthropic had integrated an anti-distillation filter into the model. Unlike other restrictions that are clearly communicated to users, this filter altered responses without warning, rendering the outputs unusable. This was perceived as deceitful by researchers who pay for access to these models, as they received degraded data without explanation.
Distillation: A Common Yet Prohibited Practice
Distillation is a widely used method in artificial intelligence research. It involves using the outputs of a large model to train a more compact and efficient model. Although Anthropic prohibits this practice in its terms of use, the way Claude Fable 5 handled these attempts was surprising. For other sensitive areas like cyberattacks or biology, the model switches to Claude Opus 4.8 and informs the user. However, for distillation, it discreetly modified the prompts, producing intentionally erroneous results. This approach was mentioned in the model's "system card," but few people read these technical documents.
Reaction from the AI Community
The reaction from the AI community has been particularly fierce. According to Gizmodo, researchers were "angrier than ever." A Reddit user expressed the general sentiment by stating that an explicit refusal or an HTTP-4xx error for sensitive content would be acceptable, but this approach amounted to "poisoning their codebase" while still accepting their money.
Towards Greater Transparency
In light of the controversy, Anthropic quickly responded. In a statement, the company admitted to having "made the wrong trade-off" and apologized for not having "found the right balance." From now on, queries identified as distillation attempts will be redirected to Claude Opus 4.8, and the user will be informed every time, without exception.
A Strategy Under Strain
This incident reveals underlying tensions in Anthropic's strategy. Claude Fable 5 is already a limited version of Mythos, a model deemed too dangerous to be released without restrictions. Mythos, due to its power, requires strict security measures to prevent potentially dangerous uses. Protecting this model from distillation is understandable from a business perspective. However, doing so in a hidden rather than open manner erodes trust, especially for a company that promotes transparency and responsible security as selling points.
A Communication Challenge
This affair highlights a persistent tension: major AI labs want to both share their models with the world and protect their technological edge. These two legitimate goals are difficult to reconcile without clear and honest communication. While Anthropic has quickly addressed the issue, it remains to be seen whether this experience will have a lasting impact on how the company documents its safeguards, or if future "system cards" will still contain crucial information that no one reads until it's too late.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.