用道德心理学理论分析大模型的伦理判断差异
Differences in the Moral Foundations of Large Language Models
- 基于哈特道德基础理论,测试多款主流大模型的价值判断
- 模型越强大,其道德倾向与人类基准差异越大
- 揭示模型间道德立场分歧,提醒政策制定者关注对齐问题
大语言模型在政治、商业和教育等关键领域日益广泛应用,但其规范性伦理判断的本质仍不清晰。现有对齐研究尚未充分借鉴道德心理学的视角来指导前沿模型的训练与评估。本文基于乔纳森·哈特的道德基础理论(MFT),对多家主流模型提供商的多种模型进行合成实验,以激发模型产生多样化的价值判断。通过多种描述性统计方法,记录了模型响应相对于原始调查中人类基准的偏差与方差。结果表明,不同模型之间以及模型与全国代表性人类基准之间,在道德基础依赖上存在显著差异,且随着模型能力提升,这种差异愈发明显。该研究旨在推动未来使用MFT对大模型进行更深入分析,包括开源模型的微调,并促使政策制定者更加重视道德基础在大模型对齐中的重要性。
原文摘要 · Abstract (English)
Large language models are increasingly being used in critical domains of politics, business, and education, but the nature of their normative ethical judgment remains opaque. Alignment research has, to date, not sufficiently utilized perspectives and insights from the field of moral psychology to inform training and evaluation of frontier models. I perform a synthetic experiment on a wide range of models from most major model providers using Jonathan Haidt's influential moral foundations theory (MFT) to elicit diverse value judgments from LLMs. Using multiple descriptive statistical approaches, I document the bias and variance of large language model responses relative to a human baseline in the original survey. My results suggest that models rely on different moral foundations from one another and from a nationally representative human baseline, and these differences increase as model capabilities increase. This work seeks to spur further analysis of LLMs using MFT, including finetuning of open-source models, and greater deliberation by policymakers on the importance of moral foundations for LLM alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。