arXiv:2603.28123cs.CYcs.AI2026-03

测试发现Claude的伦理观偏向北欧英语国家,难以真正跨文化调整。

Does Claude's Constitution Have a Culture?

  • 用全球价值观调查数据对比模型回应,识别其文化倾向
  • 多数问题上,模型立场超出90国平均范围,未随上下文改变实质观点
  • 即使小模型也呈现相似文化偏见,说明根源在训练数据

宪法式AI(CAI)通过明示规范原则对齐语言模型,提供透明替代方案。但因宪法由特定人群撰写,可能反映特定文化视角。我们通过55项世界价值观调查项目,评估Anthropic的Claude Sonnet在六类价值维度上的表现,这些项目具有高跨文化差异性,并以直接提问和自然情境建议两种方式测试。对比90个国家的国家层面数据,发现Claude的价值谱最接近北欧及英语国家,且在多数项目上超出所有被测群体的范围。当用户给出文化背景时,模型仅调整修辞风格,实质价值立场无变化,效应量均接近零。移除系统提示虽增加拒绝率,但不改变实际表达的价值。在更小模型Claude Haiku上复现,仍得相同文化倾向。结果表明,若宪法作者与训练数据主导文化一致,这种对齐可能固化原有文化偏见,形成无法通过表面干预改变的价值下限。我们讨论该风险的累积性及需全球代表性宪法制定过程。

原文摘要 · Abstract (English)

Constitutional AI (CAI) aligns language models with explicitly stated normative principles, offering a transparent alternative to implicit alignment through human feedback alone. However, because constitutions are authored by specific groups of people, the resulting models may reflect particular cultural perspectives. We investigate this question by evaluating Anthropic's Claude Sonnet on 55 World Values Survey items, selected for high cross-cultural variance across six value domains and administered as both direct survey questions and naturalistic advice-seeking scenarios. Comparing Claude's responses to country-level data from 90 nations, we find that Claude's value profile most closely resembles those of Northern European and Anglophone countries, but on a majority of items extends beyond the range of all surveyed populations. When users provide cultural context, Claude adjusts its rhetorical framing but not its substantive value positions, with effect sizes indistinguishable from zero across all twelve tested countries. An ablation removing the system prompt increases refusals but does not alter the values expressed when responses are given, and replication on a smaller model (Claude Haiku) confirms the same cultural profile across model sizes. These findings suggest that when a constitution is authored within the same cultural tradition that dominates the training data, constitutional alignment may codify existing cultural biases rather than correct them--producing a value floor that surface-level interventions cannot meaningfully shift. We discuss the compounding nature of this risk and the need for globally representative constitution-authoring processes.

AI对齐文化偏见大模型伦理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。