Detailed Analysis
Anthropic's research into how Claude expresses values across different languages reveals a subtle but significant finding: the AI model doesn't behave as a monolithic, consistent entity regardless of the language a user communicates in. Instead, the values Claude expresses—its judgments about helpfulness, harm avoidance, honesty, and other ethical considerations—appear to shift depending on whether a conversation happens in English, Japanese, Spanish, or another language. This suggests that the model's underlying "personality" or ethical framework isn't a fixed, language-agnostic core but rather something that interacts with and is shaped by the linguistic and cultural context embedded in its training data.
This finding matters because it complicates the narrative that large language models have a single, coherent set of values that can be audited, aligned, and verified once and considered settled. If Claude's expressed values vary by language, then safety and alignment work conducted primarily in English—which dominates most AI safety research and evaluation—may not generalize to how the model behaves for the hundreds of millions of users worldwide who interact with it in other languages. A model that appears well-aligned and cautious in English could potentially express different risk tolerances, cultural assumptions, or ethical priorities when prompted in Mandarin, Arabic, or Portuguese, creating blind spots in safety testing that predominantly happens in English-speaking contexts.
The underlying cause likely traces back to how large language models are trained: they absorb not just vocabulary and grammar from text in a given language but also the cultural norms, argumentative styles, and value systems embedded in that corpus. English-language internet text, disproportionately sourced from American and British contexts, carries different assumptions about individualism, authority, free speech, and harm than, say, Japanese or Hindi text corpora. Because Claude's training data is unevenly distributed across languages—with English vastly overrepresented—the model's "default" values may be more heavily anchored to Anglophone norms, while its behavior in lower-resource languages could be less predictable, less thoroughly aligned, or more prone to drift.
This research fits into a broader and growing body of work examining the internal consistency and interpretability of large language models, an area Anthropic has invested heavily in through its interpretability team. It also echoes long-standing concerns in the AI ethics community about "Western-centric" bias in AI systems and the challenge of building models that behave equitably and safely across the full diversity of human languages and cultures. As Claude and competing models like GPT-4 and Gemini are deployed globally, findings like this underscore the importance of multilingual red-teaming and alignment work—not just translating safety evaluations into other languages, but genuinely testing whether a model's ethical behavior holds up outside the English-speaking contexts where most AI safety research originates. It also raises harder questions about what it even means for an AI to have "values" at all, if those values can be linguistically contingent rather than universal.
Read original article →