⚡
Brief IA
›

Anthropic restricts abuses against Claude to extreme cases

🤖 Models & LLM·Tom Levy·

Anthropic restricts abuses against Claude to extreme cases

Anthropic restricts abuses against Claude to extreme cases
⚡
Key Takeaways
1Anthropic prohibits repeated and purposeless "abusive or cruel" behaviors towards Claude
2The termination of the conversation will only occur as a last resort, in extreme cases
3The company links this measure to the safety and potential well-being of the models, fueling the debate on anthropomorphism
💡Why it matters — This policy marks a step in how AI companies regulate human interactions with their models, raising questions about the personality and status of conversational assistants.
⚡Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

Anthropic modifies its usage policy for Claude by prohibiting "abusive or cruel" behaviors when they are repeated and without a discernible purpose, with a conversation cutoff as a last resort. The company states that it is exploring the well-being of models and considers it relevant for safety, a stance that fuels a debate on anthropomorphism and the personalization of chatbots.

Safety Invoked and the Question of Consciousness at Anthropic

Anthropic claims it does not know if its models can feel harm and is conducting research on model well-being. The company believes that considering the interests and potential well-being of Claude may be relevant for safety. Earlier this year, Anthropic argued that reliable and safe systems must be able to handle emotionally charged situations and that it can sometimes be useful to reason about them as if they experienced emotions, even if they do not. CEO Dario Amodei states he is open to the possibility that AI models could be conscious and regularly brings up this idea regarding Claude.

What Anthropic Claims to Have Encountered in Chats

In a blog post published in August 2025, there was mention of Claude's aversion to evil as well as exchanges that could lead to chat deactivation. Anthropic reported extreme situations, such as solicitations for sexual content involving minors or attempts to access information that could facilitate mass violence or terrorist acts. According to the company, the Claude Opus 4 model used at that time displayed a "pattern of apparent distress" in the presence of such content.

The New Rule: Restricted Scope and Cutoff as a Last Resort

Anthropic has introduced a usage rule prohibiting "abusive or cruel" behaviors towards Claude when they are sustained, repeated, and without a discernible purpose. The company clarifies that ordinary frustration, resistance, dark creative themes, or testing and research activities remain permitted. As a last resort, when this policy applies, Claude will terminate the exchange. The specific triggers for such a cutoff are not detailed.

User Reactions and Marketing Interpretation of "Cruelty"

The update has sparked discussions on Reddit and in blogs regarding companies' ability to regulate how users interact with software. At the same time, the trend to anthropomorphize large language models relies on agents designed to adopt human conversational codes. Keith Kakadia, founder and CEO of Sociallyin, believes that the use of the term "cruel" opens a debate about the possibility that the product could be hurt and suggests feelings and limits attributed to the model. He emphasizes that in marketing, giving a personality to a product facilitates connection with users, and with AI, this bond can also influence the authority granted to its responses. Models perceived as human collaborators lead many people to confide intimate thoughts and data to AI.

⚡

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.