arXiv:2507.17216cs.CLcs.AI2025-07Conference of the …被引 10

发现大模型与人类在道德判断上存在显著差异,尤其在意见分歧时。

The Pluralistic Moral Gap: Understanding Judgment and Value Differences between Humans and Large Language Models

  • 构建1618个真实道德困境数据集,对比人类与模型判断分布
  • 模型仅在人类共识高时对齐,分歧时表现大幅下降
  • 提出动态道德画像方法,使模型输出更贴近人类多元价值

随着人们越来越多依赖大语言模型(LLMs)获取道德建议,其对人类决策的影响日益显著。然而,当前对模型与人类道德判断的一致性了解甚少。为此,我们构建了包含1,618个真实道德困境的道德困境数据集,每个困境配有基于人类判断的分布,包括二元评价和自由文本理由。我们将此问题视为一种多元分布对齐任务,比较模型与人类在不同困境下的判断分布。结果发现,模型仅在人类高度一致时能复现人类判断;当人类分歧增加时,对齐度急剧下降。同时,通过从3,783条理由中提取的60种价值分类体系,我们发现模型所依赖的道德价值范围远窄于人类。这些结果揭示了一个‘多元道德差距’:不仅判断分布不匹配,且价值多样性不足。为缩小这一差距,我们提出动态道德画像(DMP),一种基于狄利克雷采样的条件生成方法,以人类价值谱为依据。DMP使对齐度提升64.3%,并显著增强价值多样性,为实现更包容、更贴近人类的道德指导迈出关键一步。

原文摘要 · Abstract (English)

People increasingly rely on Large Language Models (LLMs) for moral advice, which may influence humans' decisions. Yet, little is known about how closely LLMs align with human moral judgments. To address this, we introduce the Moral Dilemma Dataset, a benchmark of 1,618 real-world moral dilemmas paired with a distribution of human moral judgments consisting of a binary evaluation and a free-text rationale. We treat this problem as a pluralistic distributional alignment task, comparing the distributions of LLM and human judgments across dilemmas. We find that models reproduce human judgments only under high consensus; alignment deteriorates sharply when human disagreement increases. In parallel, using a 60-value taxonomy built from 3,783 value expressions extracted from rationales, we show that LLMs rely on a narrower set of moral values than humans. These findings reveal a pluralistic moral gap: a mismatch in both the distribution and diversity of values expressed. To close this gap, we introduce Dynamic Moral Profiling (DMP), a Dirichlet-based sampling method that conditions model outputs on human-derived value profiles. DMP improves alignment by 64.3% and enhances value diversity, offering a step toward more pluralistic and human-aligned moral guidance from LLMs.

道德判断大模型对齐价值多样性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。