Anthropic to Ban Abuse Against Claude

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Anthropic plans to classify sustained and deemed unnecessary abuse towards its assistant Claude as a violation of usage. This measure, presented as a standard of responsibility, raises questions about its scope, application, and potential effects on model safety. Executives and investors are divided on the necessity, ethics, and effectiveness of such a rule.
Scope and Application Contested by Investors and Founders
Investors and entrepreneurs are questioning the concrete scope of a rule governing user tone. Mark Verner points out that AI companies already control user-generated content and could now regulate how users address the models. He wonders, for example, about asking Claude which product could replace it and whether this could be perceived as abuse, arguing that it will depend on how Anthropic interprets and applies the term "unnecessary." Lachlan Phillips suggests focusing on product design rather than sanctioning accounts to respond to hostile user behavior. He proposes training Claude not to react and to complete the requested task, or filtering abusive messages to transform them into standard queries. According to him, the rule imposes a worldview on users more than it actually protects the model.
The Rule and Its Official Justification Do Not Explicitly Target Training
Anthropic states that subjecting Claude to repeated and unjustified abuse will be considered a violation of its usage policy. The company describes this initiative as a standard aimed at preventing harm and encouraging responsible use, without making a direct link to model training. This lack of a direct connection leads Bill Gurley to question the use of customer prompts for training models and to assert that if this were not the case, the rule would have no impact on the system as a whole.
Arguments for More Courteous Interactions for Safety
Some argue that this approach could contribute to the safety and alignment of systems. Aaron Levie, co-founder and CEO of Box, finds the policy unusual but potentially prudent and clarifies that he does not consider AI to be conscious. He emphasizes the importance of how models are exposed to interactions during their training and use, believing it makes sense to avoid future models being trained on content where humans are rude to them. He describes politeness towards AI as an "easy Pascal's wager," suggesting that more positive exchanges could foster safer and better-aligned systems.
Anthropomorphism and Consciousness: Critics Reject the Assimilation to Humans
Several voices criticize the anthropomorphism that such a policy could induce. Michael Shellenberger believes that Anthropic treats machines too similarly to humans, reminding that the company has already expressed doubts about Claude's potential moral status and its desire to shield it from stressful exchanges. Scott Stevenson mentions a risky precedent, asserting that there is no reason to assume that current models possess consciousness and that their ability to generate art, philosophy, or software does not constitute proof of consciousness. Joseph Carlson, who claims to have never been cruel to a model but finds the rule very strange, states that Claude has no emotions, moral responsibility, or consciousness. He believes that labeling abusive prompts as harmful to the model risks confusing users' perceptions of the nature of AI and likens rudeness towards Claude to that which one might show towards a piece of furniture.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.