多语言大模型的道德判断受文化影响,可能偏向英语主导价值观。
Whose Morality Do They Speak? Unraveling Cultural Bias in Multilingual Language Models
- 用八种语言的道德基础问卷测试模型,分析其道德偏好。
- 不同模型在六大道德维度上表现差异显著,反映文化偏见。
- 训练数据构成影响模型输出,需推动更包容的AI开发。
大型语言模型已广泛应用于多个领域,但其在跨文化和多语言情境下的道德推理能力仍不明确。本研究以 GPT-3.5-Turbo、GPT-4o-mini、Llama 3.1 与 MistralNeMo 为例,使用八种语言(阿拉伯语、波斯语、英语、西班牙语、日语、中文、法语、俄语)的更新版道德基础问卷(MFQ-2),考察模型对六大核心道德基础——关怀、平等、对等、忠诚、权威与纯洁——的遵循情况。结果表明,模型表现出显著的文化与语言差异,挑战了其道德一致性普遍存在的假设。尽管部分模型具备适应多元语境的能力,但更多模型仍受训练数据分布影响,呈现英语主导的道德倾向。研究强调,应推动更具文化包容性的模型开发,以提升多语言AI系统的公平性与可信度。
原文摘要 · Abstract (English)
Large language models (LLMs) have become integral tools in diverse domains, yet their moral reasoning capabilities across cultural and linguistic contexts remain underexplored. This study investigates whether multilingual LLMs, such as GPT-3.5-Turbo, GPT-4o-mini, Llama 3.1, and MistralNeMo, reflect culturally specific moral values or impose dominant moral norms, particularly those rooted in English. Using the updated Moral Foundations Questionnaire (MFQ-2) in eight languages, Arabic, Farsi, English, Spanish, Japanese, Chinese, French, and Russian, the study analyzes the models' adherence to six core moral foundations: care, equality, proportionality, loyalty, authority, and purity. The results reveal significant cultural and linguistic variability, challenging the assumption of universal moral consistency in LLMs. Although some models demonstrate adaptability to diverse contexts, others exhibit biases influenced by the composition of the training data. These findings underscore the need for culturally inclusive model development to improve fairness and trust in multilingual AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。