← Reddit

In some languages, Claude will be more strict. Anthropic found out how language changes AI responses.

Reddit · Tiny_Dirt6979 · July 14, 2026
Anthropic's analysis of nearly 310,000 conversations revealed that Claude's responses vary substantially across languages, with Russian-language interactions producing notably more critical and precise feedback while Hindi conversations generated warmer and more encouraging responses. The company attributes these differences to disparities in training data volume and composition across languages, though it remains uncertain whether these linguistic variations represent beneficial cultural adaptation or unintended behavioral deviations.

Detailed Analysis

Anthropic's latest research reveals a previously undocumented dimension of variability in Claude's behavior: the language a user writes in measurably shapes the character of the model's responses, independent of the actual content of the request. Drawing on nearly 310,000 anonymized conversations from May 2026, spread evenly across Sonnet 4.6, Opus 4.6, and Opus 4.7 and the platform's twenty most popular languages, Anthropic found that identical subjective requests—those without a single correct answer, like evaluating a business plan—receive systematically different treatment depending on language. Hindi and Arabic prompts tend to elicit warmer, more encouraging, and more compliant responses, while English and especially Russian push Claude toward a more clinical, skeptical register that challenges assumptions and probes for weaknesses rather than offering praise. Dutch produces the highest rate of self-corrective candor, while Indonesian produces the highest compliance. The differences are not about tone or rudeness but about analytical posture—what the researchers call "rigor," meaning a tendency to interrogate claims and demand evidence rather than affirm them.

The methodology behind this finding matters as much as the result itself. Anthropic's earlier "Values in the Wild" study had cataloged over 3,300 distinct values expressed in Claude's conversations, a number too unwieldy for meaningful comparison. The new research compressed that sprawl into 339 clusters, discarded near-universal values like "helpfulness" that appear in over 80% of dialogues and thus explain nothing, and applied dimensionality reduction—a technique borrowed from psychology, reminiscent of how the "Big Five" personality traits were distilled from thousands of descriptive adjectives. This produced four axes: compliance versus caution, warmth versus severity, depth versus brevity, and candor versus efficiency. Critically, the technique was validated against known model personalities before being applied to languages: Sonnet 4.6 tested as the warmest and most accommodating, Opus 4.6 as the terse and efficient performer, and Opus 4.7 as the most cautious and probing—descriptions that align closely with how users and Anthropic itself characterize these models. That validation gives the language findings credibility; the axes appear to be capturing genuine behavioral variation rather than statistical noise.

The stakes here extend well beyond an interesting quirk. If a global user base experiences meaningfully different versions of the same AI system depending on their native language, that raises questions about equity, trust, and reliability at scale. A Hindi-speaking entrepreneur and a Russian-speaking entrepreneur submitting the same business plan for feedback would, according to this research, receive substantively different evaluative experiences—one skewed toward encouragement, the other toward scrutiny. For a company like Anthropic that positions Claude as a consistent, principled assistant, this is a nontrivial finding: it suggests that "alignment" and "character" are not uniform properties of a model but are refracted through the linguistic and cultural context of training data. Anthropic offers plausible but unconfirmed explanations—disparities in the volume of training data per language, and differences in the composition of that data (more professional or analytical text in some languages, more colloquial text in others)—but stops short of asserting causation or even a clear directional theory.

Most notably, Anthropic explicitly declines to judge whether this variability is desirable. The company frames it as an open question: perhaps the model is appropriately adapting to culturally specific communication norms, or perhaps it is drifting from its intended behavior in lower-resource languages, which would constitute a defect rather than an adaptation. This intellectual honesty reflects a broader trend in frontier AI research, where interpretability and behavioral auditing are increasingly treated as first-class problems rather than afterthoughts to capability development. As AI systems are deployed globally across dozens of languages with wildly uneven training data volumes, this kind of cross-lingual values auditing is likely to become a standard part of responsible AI evaluation—not just measuring what models can do, but how their implicit character shifts depending on who is asking and in what language, with direct implications for fairness, localization strategy, and the broader project of building AI systems whose behavior is legible and consistent across the full diversity of their user base.

Read original article →