arXiv:2506.12433cs.CLcs.AI2025-06被引 2

用大模型检测不同文化道德观差异,发现越先进的模型越贴近真实人类判断。

Exploring Cultural Variations in Moral Judgments with Large Language Models

  • 通过道德合理性得分对比模型输出与全球调查数据
  • 先进指令微调模型与人类道德判断相关性显著提升
  • 模型更贴近西方国家观点,跨文化敏感性仍有不足

大型语言模型(LLMs)在多项任务中表现优异,但其对多元文化道德价值观的捕捉能力仍不明确。本文考察了多个模型是否反映世界价值调查(WVS)和皮尤研究中心全球态度调查(PEW)报告的文化道德差异。我们比较了较小的单语与多语言模型(GPT-2、OPT、BLOOMZ、Qwen)以及近期指令微调模型(GPT-4o、GPT-4o-mini、Gemma-2-9b-it、Llama-3.3-70B-Instruct)。基于对数概率的道德合理性评分,将各模型输出与涵盖广泛伦理议题的调查数据进行相关性分析。结果表明,早期或小型模型常与人类判断呈现接近零或负相关;而先进指令微调模型则展现出显著更高的正相关性,说明其更贴近现实世界道德态度。区域分析显示,模型与西方、受教育、工业化、富裕、民主(W.E.I.R.D.)国家的匹配度更高。尽管模型规模扩大和指令微调有助于提升跨文化道德一致性,但在某些主题和地区仍存在挑战。研究讨论了偏见分析、训练数据多样性、信息检索影响及提升模型文化敏感性的策略。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have shown strong performance across many tasks, but their ability to capture culturally diverse moral values remains unclear. In this paper, we examine whether LLMs mirror variations in moral attitudes reported by the World Values Survey (WVS) and the Pew Research Center's Global Attitudes Survey (PEW). We compare smaller monolingual and multilingual models (GPT-2, OPT, BLOOMZ, and Qwen) with recent instruction-tuned models (GPT-4o, GPT-4o-mini, Gemma-2-9b-it, and Llama-3.3-70B-Instruct). Using log-probability-based \emph{moral justifiability} scores, we correlate each model's outputs with survey data covering a broad set of ethical topics. Our results show that many earlier or smaller models often produce near-zero or negative correlations with human judgments. In contrast, advanced instruction-tuned models achieve substantially higher positive correlations, suggesting they better reflect real-world moral attitudes. We provide a detailed regional analysis revealing that models align better with Western, Educated, Industrialized, Rich, and Democratic (W.E.I.R.D.) nations than with other regions. While scaling model size and using instruction tuning improves alignment with cross-cultural moral norms, challenges remain for certain topics and regions. We discuss these findings in relation to bias analysis, training data diversity, information retrieval implications, and strategies for improving the cultural sensitivity of LLMs.

道德判断文化差异大模型评估跨文化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。