LLM道德判断缺乏文化多样性,可能扭曲人类价值观。
Cultural Bias in Large Language Models: Evaluating AI Agents through Moral Questionnaires
- 用19种文化道德问卷测试大模型,对比人类基准
- 模型普遍弱化文化差异,大小提升不改善表现
- 适合关注AI伦理、跨文化研究的研究者
大语言模型是否真正代表人类价值观,还是仅在平均化?本研究发现:尽管具备语言能力,大模型未能体现多元文化道德框架。我们通过在19个文化背景下应用道德基础问卷(Moral Foundations Questionnaire),揭示了模型生成与人类道德直觉间的显著差距。对比多个前沿大模型与人类基准数据,发现这些模型系统性地同质化道德多样性。令人意外的是,模型规模增大并未持续提升文化代表性。研究挑战了将大模型作为社会科学研究中合成人群的使用,并凸显当前对齐方法的根本局限——缺乏数据驱动的对齐策略,无法捕捉细微的文化特定道德直觉。结果呼吁更扎实的对齐目标与评估指标,以确保人工智能体现多样人类价值观,而非扁平化道德图景。
原文摘要 · Abstract (English)
Are AI systems truly representing human values, or merely averaging across them? Our study suggests a concerning reality: Large Language Models (LLMs) fail to represent diverse cultural moral frameworks despite their linguistic capabilities. We expose significant gaps between AI-generated and human moral intuitions by applying the Moral Foundations Questionnaire across 19 cultural contexts. Comparing multiple state-of-the-art LLMs' origins against human baseline data, we find these models systematically homogenize moral diversity. Surprisingly, increased model size doesn't consistently improve cultural representation fidelity. Our findings challenge the growing use of LLMs as synthetic populations in social science research and highlight a fundamental limitation in current AI alignment approaches. Without data-driven alignment beyond prompting, these systems cannot capture the nuanced, culturally-specific moral intuitions. Our results call for more grounded alignment objectives and evaluation metrics to ensure AI systems represent diverse human values rather than flattening the moral landscape.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。