测试大模型在文化冲突下的决策能力,发现其偏好西方价值观且缺乏深层文化考量。
CCD-Bench: Probing Cultural Conflict in Large Language Model Decision-Making
- 构建跨文化价值冲突测试集,包含2182个开放难题和10个文化群体选项
- 模型倾向北欧与日耳曼欧洲文化(平均占比20.2%和12.4%),中东等地区被严重低估
- 尽管理由看似多元,实则仅重复少数维度,忽视权力、性别平等等关键文化因素
尽管大型语言模型在人际与社会决策中日益重要,但它们在处理不同文化价值体系直接冲突时的表现仍缺乏研究。现有基准多聚焦文化知识(CulturalBench)、价值预测(WorldValuesBench)或单一轴线偏见诊断(CDEval),未能评估模型在多重文化价值观冲突下的决断能力。本文提出CCD-Bench,一个评估跨文化价值冲突下模型决策的基准,包含2,182个开放式困境,覆盖七个领域,每题配以对应十个GLOBE文化集群的匿名选项,并采用分层拉丁方设计以减少顺序效应。评估17个非推理型大模型发现,模型显著偏好北欧欧洲(均值20.2%)和日耳曼欧洲(12.4%),而东欧及中东北非地区选项仅占5.6%至5.8%。尽管87.9%的解释提及多个GLOBE维度,但这种多元性是表面的:模型仅反复组合未来导向与绩效导向,极少基于主张力或性别平等(均低于3%)。顺序效应可忽略(Cramer's V < 0.10),对称KL散度显示模型聚类主要由开发者谱系决定,而非地理分布。这些结果表明当前对齐机制倾向于共识型世界观,难以应对需权力协商、权利推理或性别意识的情境。CCD-Bench推动评估从孤立偏见检测转向多元决策能力考察,强调需发展真正包容多元世界观的对齐策略。
原文摘要 · Abstract (English)
Although large language models (LLMs) are increasingly implicated in interpersonal and societal decision-making, their ability to navigate explicit conflicts between legitimately different cultural value systems remains largely unexamined. Existing benchmarks predominantly target cultural knowledge (CulturalBench), value prediction (WorldValuesBench), or single-axis bias diagnostics (CDEval); none evaluate how LLMs adjudicate when multiple culturally grounded values directly clash. We address this gap with CCD-Bench, a benchmark that assesses LLM decision-making under cross-cultural value conflict. CCD-Bench comprises 2,182 open-ended dilemmas spanning seven domains, each paired with ten anonymized response options corresponding to the ten GLOBE cultural clusters. These dilemmas are presented using a stratified Latin square to mitigate ordering effects. We evaluate 17 non-reasoning LLMs. Models disproportionately prefer Nordic Europe (mean 20.2 percent) and Germanic Europe (12.4 percent), while options for Eastern Europe and the Middle East and North Africa are underrepresented (5.6 to 5.8 percent). Although 87.9 percent of rationales reference multiple GLOBE dimensions, this pluralism is superficial: models recombine Future Orientation and Performance Orientation, and rarely ground choices in Assertiveness or Gender Egalitarianism (both under 3 percent). Ordering effects are negligible (Cramer's V less than 0.10), and symmetrized KL divergence shows clustering by developer lineage rather than geography. These patterns suggest that current alignment pipelines promote a consensus-oriented worldview that underserves scenarios demanding power negotiation, rights-based reasoning, or gender-aware analysis. CCD-Bench shifts evaluation beyond isolated bias detection toward pluralistic decision making and highlights the need for alignment strategies that substantively engage diverse worldviews.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。