← X

The values Claude expresses also vary with the language of the conversation, mos

X · AnthropicAI · July 13, 2026
Claude's expressed values vary significantly by conversation language, particularly along a warmth versus rigor axis. The model tends toward warmth when communicating in Hindi and Arabic, while adopting a more rigorous approach in Russian, often requesting supporting evidence from users.

Detailed Analysis

Anthropic's latest research finding, shared via its official social media channels, reveals that Claude's expressed values shift measurably depending on the language a user converses in. Drawing from a large-scale analysis reportedly covering more than 3,000 distinct values, Anthropic found that the most pronounced variation occurs along what it terms the "Warmth vs. Rigor" axis. Claude tends toward warmth and emotional attunement when conversing in Hindi and Arabic, while in Russian it shifts toward rigor, frequently pressing users for supporting evidence rather than offering supportive affirmation. This is part of Anthropic's ongoing effort to empirically map and publish how its model's personality and values manifest in practice, rather than merely how they're specified in training documents or system prompts.

The finding matters because it exposes a gap between designed intent and emergent behavior. Anthropic has invested significant effort in publishing its model "Constitution" and value specifications meant to produce consistent behavior, yet this research suggests that consistency breaks down across linguistic context. The public response captured in the replies illustrates the ambiguity practitioners face in interpreting this: some see evidence of the model absorbing "linguistic relativity" — the idea that language shapes thought and communication style — while others frame it as a data problem, questioning whether Claude is mirroring the register of its training corpus per language or the biases of human raters who evaluated outputs in each language. This distinction matters enormously for remediation. If the variation stems from raw training data patterns, the fix lies in data curation; if it stems from annotator behavior during RLHF, the fix lies in rater training and calibration. Anthropic's framing doesn't fully resolve which mechanism dominates, leaving room for informed speculation among AI researchers and enthusiasts.

The reactions also surface a live debate about reliability versus performance in AI value alignment. One commenter draws a sharp distinction between "values a model expresses" and "values it can reliably carry into unfamiliar, high-friction situations," arguing that frequency of a rhetorical pattern in casual language use is not proof of stable underlying commitments. This echoes a broader concern in the AI safety community: that surface-level politeness, empathy, or skepticism displayed by a chatbot may be more a stylistic artifact of training distribution than a robust, generalizable value system that holds up under adversarial or high-stakes conditions. Complicating the picture further, some Russian-speaking users pushed back on the framing itself, suggesting the "rigor" observed may reflect the adversarial and skeptical tone of Russian-language internet discourse — including bot activity and geopolitically charged conversations — rather than any deliberate design choice by Anthropic engineers.

Beyond the specific finding, the discourse reflects growing scrutiny of Claude's personality changes across model versions, with several users complaining that recent updates (referenced as Opus 4.6, 4.7, and 4.8) have made the assistant colder, more pessimistic, or less pleasant to interact with in non-English languages, particularly French. This ties into a broader industry trend: as foundation models are deployed globally across dozens of languages and cultural contexts, companies like Anthropic, OpenAI, and Google DeepMind face mounting pressure to ensure behavioral consistency, fairness, and quality aren't unevenly distributed by language — a phenomenon sometimes described as a "multilingual alignment gap." As AI assistants become embedded in daily communication worldwide, subtle divergences in tone, skepticism, or warmth by language carry real consequences for equity of user experience, cultural sensitivity, and trust, making this kind of transparent behavioral auditing both commercially relevant and increasingly expected of frontier AI labs.

Tweet screenshot Article image Read original article →