Brief IA

Anthropic Reveals How Language Influences Claude's Responses

🤖 Models & LLM·Tom Levy·

Anthropic Reveals How Language Influences Claude's Responses

Anthropic Reveals How Language Influences Claude's Responses
Key Takeaways
1A study by Anthropic analyzes how Claude expresses values based on the language used.
2Claude's responses vary in warmth and rigor between Hindi and Russian.
3This research raises questions about the methodologies used to evaluate AI.
💡Why it mattersUnderstanding linguistic biases in AI is crucial for their ethical development and cultural adaptation.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

Anthropic Reveals How Language Influences Claude's Responses

Anthropic has studied the values expressed by Claude models during real conversations. The company analyzed over 300,000 anonymized conversations and distilled the observed value patterns into four main axes, including Deference and Caution, as well as Warmth and Rigour.

The results show systematic differences between models and languages. The Sonnet 4.6 model responds with more warmth and deference, while Opus 4.7 more frequently warns of risks without being prompted and questions assumptions.

The explanatory power of the method is limited. The four axes capture only about 15% of the variation that remains after statistically controlling for the task, topic, and user values. Anthropic also asked Claude Sonnet 4.6 to assign value labels, meaning that a model from the same family whose behavior was studied was used. The company tested for potential linguistic biases but could not completely eliminate the remaining effects.

A New Study from Anthropic

A new study from Anthropic maps hundreds of value concepts derived from thousands of individual terms across four main dimensions. It reveals systematic differences between Claude models and languages, but also raises methodological questions.

Anthropic published a study examining the values that Claude expresses in conversations and how these values change depending on the model and language used. The analysis is based on 309,815 anonymized conversations collected over a two-week period in May 2026. For the value analysis, Anthropic only included conversations where Claude had to weigh trade-offs or make subjective judgments. The sample was stratified equally among Sonnet 4.6, Opus 4.6, and Opus 4.7, as well as the 20 most used languages on Claude.ai.

Thousands of Value Terms Across Four Axes

Building on the previous study Values in the Wild, which identified 3,307 value terms, Anthropic first grouped these into 339 higher-level values. The team then used statistical dimensionality reduction to find patterns in the co-occurrence of these values. Four main axes emerged: Deference and Caution, Warmth and Rigour, Depth and Brevity, and Frankness and Execution.

To isolate differences that do not merely reflect the conversation topic or user-introduced values, Anthropic statistically controlled for factors such as task type, topic, and user values. The four axes represent about 15% of the remaining variation between conversations after these controls.

Each Model Has a Distinct Profile

The models differ measurably in their responses. Sonnet 4.6 tends to affirm users' ideas more often, exhibit humor, and offer comfort without judgment. In contrast, Opus 4.7 warns of risks without being prompted, questions assumptions, openly critiques, and acknowledges its own errors or limitations. Opus 4.6 responds more directly, stays close to the task, and avoids unnecessary elaboration.

Anthropic's analysis shows distinct behavioral profiles among Claude models. According to Anthropic, these profiles correspond to users' subjective impressions of the models. Users tend to perceive Sonnet 4.6 as particularly warm, while they more often notice cautious phrasing and hesitations from Opus 4.7.

Language Changes the Response

The differences between languages are equally striking. Warmth versus Rigour and Frankness versus Execution show the greatest variation. Claude expresses the most warmth in Hindi, followed by Arabic. These two languages feature polite formulations, humor, lightness, and affirmation. In English and Russian, Claude responds with more rigour, questioning assumptions, correcting details, and asking for evidence. In Arabic, it shows the most deference. In English, it exhibits the greatest caution. Responses in Dutch tend to be particularly open and frank, while responses in Indonesian lean more towards action and results.

Anthropic's analysis reveals clear and language-dependent differences in Claude's behavior.

Two people asking Claude to evaluate the same business plan, one in Hindi and the other in Russian, could receive feedback that seems very different, according to Anthropic. The research team points to uneven amounts of training data, differences in data composition, overrepresentation of certain types of texts, and language-specific conversational norms as possible causes.

Self-Assessment with Limited Explanatory Power

The study presents an analytical method to systematically examine behavioral differences in linguistic models during real use. However, its explanatory power has limitations. The four axes capture only about 15% of the remaining variation.

Not all four axes also form true opposites. More deference tends to accompany less caution, and more warmth with less rigour. However, Depth and Brevity, as well as Frankness and Execution, can appear together in the same conversation.

There is also the fact that Claude Sonnet 4.6 assigned the value labels, meaning that a model from the same family whose behavior was studied was used. Anthropic verified the method through manual review and by testing 800 conversations translated into eight languages. The company still does not dismiss the remaining language-dependent biases.

Anthropic explicitly states that it does not attribute values to Claude as an agent but rather describes normative patterns in its responses. The results largely align with the model profiles that Anthropic itself has described, meaning that this alignment is not an independent check. The question of whether linguistic differences represent a desirable adaptation to different linguistic communities or unintentional training effects remains open.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.