Anthropic Revises Its Opaque Policy on Claude's LLMs
Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Anthropic recently announced a significant revision of its policy regarding the safeguards used in the development of LLMs (large language models) with Claude Fable 5. This decision comes after criticism concerning the lack of transparency of these security measures. In a statement to WIRED, Anthropic admitted to making a poor choice and apologized for not finding the right balance.
Anthropic's initial policy, integrated into their system card, allowed Claude Fable/Mythos to detect "requests targeting the development of cutting-edge LLMs" and to "limit their effectiveness" without informing users. This approach sparked a strong backlash from the AI research community.
Announced Changes
According to information shared by @ClaudeDevs on Twitter, Anthropic has begun rolling out changes to make these safeguards visible. From now on, flagged requests will be visibly returned to Opus 4.8, similar to the safeguards for cyber and bio. Users will be able to see these notifications whenever they occur.
Regarding the API, any flagged request will return an explanation for its denial. This feature will be available server-side in the coming days.
Anthropic explains that the initial goal was to deploy Fable 5 quickly and safely. The invisible safeguards, while more targeted, were chosen to allow for rapid delivery with few false positives. However, the company acknowledges that this approach was misguided and that increased transparency is necessary for users to understand the measures in place and their rationale.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.