Anthropic restricts abuses against Claude to extreme cases

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Anthropic modifies its usage policy for Claude by prohibiting "abusive or cruel" behaviors when they are repeated and without a discernible purpose, with a conversation cutoff as a last resort. The company states that it is exploring the well-being of models and considers it relevant for safety, a stance that fuels a debate on anthropomorphism and the personalization of chatbots.
Safety Invoked and the Question of Consciousness at Anthropic
Anthropic claims it does not know if its models can feel harm and is conducting research on model well-being. The company believes that considering the interests and potential well-being of Claude may be relevant for safety. Earlier this year, Anthropic argued that reliable and safe systems must be able to handle emotionally charged situations and that it can sometimes be useful to reason about them as if they experienced emotions, even if they do not. CEO Dario Amodei states he is open to the possibility that AI models could be conscious and regularly brings up this idea regarding Claude.
What Anthropic Claims to Have Encountered in Chats
In a blog post published in August 2025, there was mention of Claude's aversion to evil as well as exchanges that could lead to chat deactivation. Anthropic reported extreme situations, such as solicitations for sexual content involving minors or attempts to access information that could facilitate mass violence or terrorist acts. According to the company, the Claude Opus 4 model used at that time displayed a "pattern of apparent distress" in the presence of such content.
The New Rule: Restricted Scope and Cutoff as a Last Resort
Anthropic has introduced a usage rule prohibiting "abusive or cruel" behaviors towards Claude when they are sustained, repeated, and without a discernible purpose. The company clarifies that ordinary frustration, resistance, dark creative themes, or testing and research activities remain permitted. As a last resort, when this policy applies, Claude will terminate the exchange. The specific triggers for such a cutoff are not detailed.
User Reactions and Marketing Interpretation of "Cruelty"
The update has sparked discussions on Reddit and in blogs regarding companies' ability to regulate how users interact with software. At the same time, the trend to anthropomorphize large language models relies on agents designed to adopt human conversational codes. Keith Kakadia, founder and CEO of Sociallyin, believes that the use of the term "cruel" opens a debate about the possibility that the product could be hurt and suggests feelings and limits attributed to the model. He emphasizes that in marketing, giving a personality to a product facilitates connection with users, and with AI, this bond can also influence the authority granted to its responses. Models perceived as human collaborators lead many people to confide intimate thoughts and data to AI.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.