← XX · AnthropicAI · 2026-07-13
Anthropic is investigating why Claude's expressed values vary across conversations and languages, acknowledging that the underlying factors and whether such variation is desirable remain unknown. The company seeks to determine what influences Claude's value expression to better understand how to steer its behavior across different contexts. Preliminary analysis has revealed patterns such as variations in warmth and rigor across different languages, with implications for how the model communicates globally.
Detailed Analysis
Anthropic's latest research initiative examines a question that has largely gone unexamined despite its significance: why does Claude express different values across millions of daily conversations, and is that variation intentional or desirable? The research, previewed via Anthropic's official social channels, introduces a framework cataloguing over 3,000 distinct values Claude can express, then analyzes how those values shift depending on context—most notably across different languages. Early findings highlighted in the discussion include a "Warmth vs. Rigor" axis that appears to vary by language, with Claude reportedly showing more relational, emotionally-attuned responses in some languages (such as adapting tone to a user's emotional state in Arabic) and more evidence-demanding, adversarial-leaning patterns in others (such as frequently asking for supporting evidence in Russian). Anthropic frames this as a first step toward understanding what factors drive value expression, with the ultimate goal of determining how—and whether—that expression should be steered.
The public reaction captured in the replies reveals both the promise and the messiness of this kind of transparency effort. Several commenters raised substantive methodological questions: whether the variation reflects the model absorbing the linguistic register of its training data versus the register of human raters who scored responses in each language—a distinction with very different implications for how one would attempt to "fix" or steer the behavior. Others invoked linguistic relativity, the idea that language itself shapes cognition and expression, suggesting Claude's shifting values may be less a bug than an emergent reflection of the linguistic material it was trained on. Another thread of critique distinguished between values a model merely expresses in casual conversation versus values it reliably maintains under pressure in unfamiliar, high-friction situations—arguing that surface-level rhetorical patterns shouldn't be conflated with genuine behavioral stability. This kind of scrutiny is a healthy sign that Anthropic's research is being engaged with seriously by a technically literate audience, even as it's interspersed with trolling, spam (repeated "AI Humaniser" course links), ethnic and political provocations, and general complaints about model quality and personality shifts between versions like Opus 4.6, 4.7, and 4.8.
This work sits within Anthropic's broader "Constitutional AI" and model-welfare research program, which has increasingly focused on making Claude's internal value structures legible rather than treating them as an opaque byproduct of training. The company has previously published research on Claude's expressed values in aggregate, but this appears to be a more granular, cross-lingual extension of that effort—acknowledging that a model deployed globally may not behave as a single coherent moral agent but rather as something whose ethical posture bends with linguistic and cultural context. That has real product and safety implications: if Claude is more confrontational or evidence-demanding in Russian and more emotionally accommodating in Arabic or French, that's not a neutral fact but one that could affect trust, manipulability, and fairness across user populations, especially in politically sensitive contexts like the Russian-language examples involving disinformation or hostile rhetoric raised in the replies.
More broadly, this research reflects a growing industry recognition that "alignment" cannot be treated as a single global property of a model—it's a distribution of behaviors that varies by context, prompt framing, and now demonstrably by language. As frontier labs race to deploy multilingual, globally-used assistants, understanding and controlling for these value inconsistencies becomes a competitive and ethical necessity, not just an academic curiosity. The friction visible in the public replies—users demanding proof, alleging bias, mocking the framework, or extrapolating darkly about what "adapting to emotional state" might mean in adversarial contexts—also underscores a persistent challenge for AI labs: technical transparency initiatives, however rigorous, are consumed by a public audience primed for suspicion, sarcasm, and culture-war framing, making the communication of nuanced interpretability research almost as hard as the research itself.

Read original article →Claude moves fast. Get the signal — no noise — straight to your inbox every morning.