研究大模型在道德选择中的偏好与不一致性
From Stability to Inconsistency: A Study of Moral Preferences in LLMs
- 基于道德基础理论构建评测数据集,覆盖六类核心道德维度
- 顶尖模型道德偏好高度相似,但对同类问题答案不一致
- 揭示当前大模型在道德推理中存在隐性偏见与认知矛盾
随着大型语言模型(LLMs)日益融入日常生活,理解其隐含偏见和道德倾向变得至关重要。为此,我们提出了一个基于道德基础理论(Moral Foundations Theory)的LLM道德评测数据集(MFD-LLM),该理论将人类道德划分为六个核心基础。我们提出一种新评估方法,通过回答一系列真实世界道德困境,全面捕捉大模型所呈现的道德偏好。研究发现,当前最先进的模型在价值偏好上表现出显著同质性,但在具体判断中却缺乏一致性。
原文摘要 · Abstract (English)
As large language models (LLMs) increasingly integrate into our daily lives, it becomes crucial to understand their implicit biases and moral tendencies. To address this, we introduce a Moral Foundations LLM dataset (MFD-LLM) grounded in Moral Foundations Theory, which conceptualizes human morality through six core foundations. We propose a novel evaluation method that captures the full spectrum of LLMs' revealed moral preferences by answering a range of real-world moral dilemmas. Our findings reveal that state-of-the-art models have remarkably homogeneous value preferences, yet demonstrate a lack of consistency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。