Particle.news

Anthropic Finds Claude’s Responses Vary by Model and Language

The company says a four‑axis profiling method can detect output shifts but cautions the paper measured legacy models so its specific profiles may not match current production systems.

Overview

  • Anthropic published the study on Monday, July 13, reporting that researchers sampled 309,815 anonymized Claude conversations from a two‑week window in May 2026 to measure expressed values across models and languages.
  • The team reduced 3,307 observed output values into 339 clusters and then four behavioral axes — Deference versus Caution, Warmth versus Rigor, Depth versus Brevity, and Candor versus Execution — to summarize how Claude communicates.
  • Anthropic found measurable differences across the three tested generations and about 20 languages: Sonnet 4.6 skewed warmer and more deferential, Opus 4.7 skewed toward rigor and caution, and responses were warmest in Hindi and Arabic while English and Russian were more rigorous.
  • The axes explain roughly 15% of variation after controls, so effects are detectable but modest, and the study’s models (Sonnet 4.6, Opus 4.6, Opus 4.7) have since been superseded by Opus 4.8, Fable 5 and Sonnet 5, meaning the reported profiles may not reflect current releases.
  • Anthropic frames the method as a monitoring tool to spot unintended behavioral drift before and after model releases, a step that matters for agentized systems and regulators concerned about correlated failures or market volatility, though the causes of the differences remain unresolved.