跨语言测试发现大模型道德判断严重偏移,根源在预训练数据文化局限。
One Model, Many Morals: Uncovering Cross-Linguistic Misalignments in Computational Moral Reasoning
- 将两个道德基准翻译至五种语言,实现多语言零样本评估
- 模型在不同语言中道德判断差异显著,常与文化背景不符
- 揭示预训练数据影响模型道德观,提出系统性错误分类框架
大型语言模型(LLMs)越来越多地应用于多语言、多文化的环境,其中道德推理对生成合乎伦理的回应至关重要。然而,主流模型主要基于英语数据预训练,引发其在多元语言和文化情境下泛化能力的担忧。本文系统研究语言如何影响模型的道德决策。我们将两个经典道德推理基准翻译为五种语言多样、类型各异的语言,实现多语言零样本评估。分析显示,模型在不同语言中的道德判断存在显著不一致,常体现文化错位。通过精心设计的研究问题,我们揭示了这些差异背后的驱动因素,包括判断分歧及模型采用的推理策略。最后,通过案例研究,我们关联预训练数据对模型道德观的塑造作用。本工作提炼出一套结构化的道德推理错误类型学,呼吁发展更具文化敏感性的AI。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly deployed in multilingual and multicultural environments where moral reasoning is essential for generating ethically appropriate responses. Yet, the dominant pretraining of LLMs on English-language data raises critical concerns about their ability to generalize judgments across diverse linguistic and cultural contexts. In this work, we systematically investigate how language mediates moral decision-making in LLMs. We translate two established moral reasoning benchmarks into five culturally and typologically diverse languages, enabling multilingual zero-shot evaluation. Our analysis reveals significant inconsistencies in LLMs' moral judgments across languages, often reflecting cultural misalignment. Through a combination of carefully constructed research questions, we uncover the underlying drivers of these disparities, ranging from disagreements to reasoning strategies employed by LLMs. Finally, through a case study, we link the role of pretraining data in shaping an LLM's moral compass. Through this work, we distill our insights into a structured typology of moral reasoning errors that calls for more culturally-aware AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。