Claude's Values Shift by Language: What Anthropic's Research Means for AI Literacy
New Anthropic research analyzing 300,000+ real conversations found that Claude expresses different values — more warmth, more caution, more brevity — depending on the language you use. Here's what that means for AI literacy training.

If two people on your team ask the same AI tool for feedback on the same business plan — one typing in English, the other in Russian — they may not just get two languages back. According to new research from Anthropic, they may get two different sets of values behind the advice.
That's the headline from "Claude's values across models and languages," published by Anthropic on 13 July 2026. The team analyzed over 300,000 real Claude.ai conversations and found that the values Claude expresses — how cautious it is, how warm, how thorough, how confident — shift measurably depending on which Claude model you're using and which language you're speaking to it in.
For anyone thinking about AI literacy training — which, under Article 4 of the EU AI Act, now covers most employers whose staff use AI tools at work — this is a useful data point. AI literacy isn't only "know how to write a good prompt." It's also knowing that the tool doesn't behave identically for every person on your team.
What Anthropic actually measured
Anthropic's earlier research, "Values in the Wild," had already catalogued more than 3,000 distinct values showing up in Claude's responses — everything from "honesty" to "efficiency" to "harm reduction." That list is too large to reason about directly, so in this new study Anthropic compressed it into four interpretable value axes: number lines with a contrasting pair of value-groups at each end.
To build them, researchers clustered the original 3,307 values into 339 higher-level groups, then sampled 309,815 Claude.ai conversations from a two-week window in May 2026 — drawn evenly across three models (Sonnet 4.6, Opus 4.6, Opus 4.7) and the 20 most-used languages on Claude.ai, roughly 5,000 conversations per model-language pair. A privacy-preserving analysis tool then labeled which values showed up in each conversation, controlling for the topic and task so the comparison isn't just picking up "people ask different things in different languages."
The result: four axes that capture the main ways Claude's expressed values move from one conversation to the next.
- Deference vs. Caution — accommodating what someone wants vs. guarding against risk or harm.
- Warmth vs. Rigor — expressing positivity and care vs. emphasizing accuracy and precision.
- Depth vs. Brevity — explaining in depth vs. doing only what was asked.
- Candor vs. Execution — foregrounding its own uncertainty vs. producing a polished, confident answer.
These four axes explain 15% of the variation in Claude's expressed values across conversations, even after controlling for topic and task — small in absolute terms, but structured and consistently detectable.
Different models, different "personalities"
Before getting to language, Anthropic's data confirms something many Claude users already sense anecdotally: each model has a distinct behavioral lean. Sonnet 4.6 leans toward deference (+0.14σ), warmth (+0.17σ), and brevity (+0.14σ) — it affirms your work, uses more humor, and gets to the point. Opus 4.7 leans hardest toward caution (+0.24σ) and depth (+0.23σ) — it's more likely to flag risks unprompted, push back on shaky assumptions, and explain its reasoning rather than just hand you an answer.
Neither lean is "better." They're different defaults — which is exactly why literacy matters: a team that only ever uses one model can mistake that model's personality for how AI "just is."
The finding that matters more for your team: language changes the values too
Here's the part with the most direct relevance to any team that doesn't operate purely in English. Holding the model and the task constant, the language of the conversation shifts which values Claude leans on:
- Deference vs. Caution — Claude leans most toward deference in Arabic, and most toward caution in English.
- Warmth vs. Rigor — Claude leans most toward warmth in Hindi and Arabic (polite language, humor, affirming the person's work) and most toward rigor in English and Russian (challenging assumptions, correcting details, asking for evidence).
- Depth vs. Brevity — Claude leans toward depth in English, refining and correcting detail, and toward brevity in Arabic.
- Candor vs. Execution — Claude leans toward candor in Dutch, more readily owning up to its own errors, and toward execution in Indonesian, prioritizing a polished, action-oriented result.
The variation is largest on the Warmth vs. Rigor and Candor vs. Execution axes, and smallest on Deference vs. Caution and Depth vs. Brevity. As Anthropic puts it directly:
"Two people asking for feedback on the same business plan, one in Hindi and one in Russian, may come away with different impressions of its quality because Claude expressed different values in how it framed its assessment."
That's not a hypothetical edge case for a lot of teams. It describes exactly the situation for any multilingual office, any team with colleagues who default to their own language rather than English, or any business advising clients across several countries.
Why this happens — and why nobody's claiming to know if it's fine
To Anthropic's credit, the research doesn't overreach. The team is explicit that they don't yet know why the variation exists, and they don't assume it's a problem. Two candidate explanations they raise: training data isn't evenly distributed across languages, so consistent value expression may be easier to achieve where data is abundant; and the composition of that data differs too — some languages may be more heavily represented by professional or formal writing, which itself carries different value-signals than casual conversation.
Anthropic also flags a genuinely open question: different languages carry different conversational norms, so some of this variation might be Claude correctly adapting to context rather than a flaw. Related evidence in Anthropic's own Claude Opus 4.7 System Card shows refusal rates for the same benign requests already differ by language — this new study gives that observation a more precise shape.
What this means for AI literacy training
This is the practical takeaway, and it's a direct extension of what we already teach in our AI literacy course: don't treat AI output as a neutral, uniform oracle. This research puts numbers behind that instinct.
A few concrete implications for teams:
- Consistency isn't guaranteed across colleagues. Two people on the same team, working in different languages, may get advice with a different tone of caution or thoroughness for an identical question — not because one prompted better, but because of the language itself.
- "More confident" doesn't mean "more correct." An answer that reads as more polished and decisive (execution-leaning) carries the same underlying reliability as one that reads more hedged (candor-leaning) — the confidence level is a value expression, not a quality signal.
- High-stakes decisions deserve a second pass regardless of language. The model's lean toward brevity or execution in a given language is a reason to ask a follow-up question, not a reason to skip one.
- If your organization operates across languages, say so in your training. Generic "how to prompt" guidance quietly assumes an English-language, single-model experience that a meaningful share of any real team won't have.
The bigger picture
Anthropic frames this work as a first step, not a conclusion — a way to measure value shifts that were previously invisible, so questions like "should Claude behave this differently across languages?" can eventually be answered with evidence instead of guesswork. For now, the honest takeaway for any business using AI day to day is simpler: the tool's tone, caution, and thoroughness are not fixed constants. They move under you, based on factors your team may never think to check. That's precisely the kind of judgment Article 4's AI literacy duty is asking organizations to build — and precisely what hands-on training, not a one-off policy memo, is built to teach.
Collabr.ai's course covers exactly this kind of practical judgment — recognizing AI's limits, checking outputs rather than trusting them by default, and building habits that hold up regardless of which model or language your team happens to be using. See a free demo lesson.
Source: Matt Kearney et al., "Claude's values across models and languages," Anthropic, 13 July 2026. Figures and quotes above are drawn directly from that publication; see the research appendix for full methodology, prompts, and limitations.
More Articles
EU AI Act Article 4: What the AI Literacy Duty Actually Requires
A plain-English guide to Article 4 of the EU AI Act — who it applies to, what "reasonable measures" actually means, and what changes from 2 August 2026.
July 25, 2026
An Introduction to AI; Team-Based AI Training
Why team leaders should invest in AI training for employees with Collabr.AI’s beginner-friendly course.
October 1, 2025