Brief IA

Anthropic Revises Its Opaque Policy on Claude's LLMs

🛠️ AI Tools·Tom Levy·

Anthropic Revises Its Opaque Policy on Claude's LLMs

Anthropic Revises Its Opaque Policy on Claude's LLMs
Key Takeaways
1Anthropic modifies its safeguards policy to make restrictions on Claude Fable 5 visible.
2The company acknowledges that its initial policy on LLMs was a mistake and apologizes.
3Users will now see why their requests are denied, with explanations provided through the API.
💡Why it mattersThis increased transparency could enhance trust among AI researchers using Anthropic's tools.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

Anthropic recently announced a significant revision of its policy regarding the safeguards used in the development of LLMs (large language models) with Claude Fable 5. This decision comes after criticism concerning the lack of transparency of these security measures. In a statement to WIRED, Anthropic admitted to making a poor choice and apologized for not finding the right balance.

Anthropic's initial policy, integrated into their system card, allowed Claude Fable/Mythos to detect "requests targeting the development of cutting-edge LLMs" and to "limit their effectiveness" without informing users. This approach sparked a strong backlash from the AI research community.

Announced Changes

According to information shared by @ClaudeDevs on Twitter, Anthropic has begun rolling out changes to make these safeguards visible. From now on, flagged requests will be visibly returned to Opus 4.8, similar to the safeguards for cyber and bio. Users will be able to see these notifications whenever they occur.

Regarding the API, any flagged request will return an explanation for its denial. This feature will be available server-side in the coming days.

Anthropic explains that the initial goal was to deploy Fable 5 quickly and safely. The invisible safeguards, while more targeted, were chosen to allow for rapid delivery with few false positives. However, the company acknowledges that this approach was misguided and that increased transparency is necessary for users to understand the measures in place and their rationale.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.