arXiv:2606.11635cs.CYcs.AI2026-06

让大模型自己制定道德评判标准,发现其能力远超此前评估结果。

Are LLMs Bad at Moral Reasoning?

  • 让大模型生成道德判断标准,而非直接评分
  • 生成的标准与人类标准更一致,且体现道德问题的复杂性
  • 适合关注大模型伦理能力评估方法的研究者

为确保高度智能的AI系统在动态开放环境中安全运行,必须具备识别、理解并响应道德理由的能力。现有研究多对当前顶尖AI系统的道德能力持悲观态度,其中一项重要工作基于1000个由人类撰写的黄金标准评判标准,对前沿大模型进行评估,结果不尽如人意。本文认为,该MoReBench数据集可被重新利用以呈现更积极的图景:若将原本要求模型给出判断的任务,改为让模型自身生成道德分析的评判标准,其生成的规则不仅比开放式回答更贴近人类标准,且在差异处反映出道德问题本身的高维度特性,甚至揭示了人类自身在制定标准时的不一致性。综合来看,该数据集表明大模型的道德推理能力远超先前估计。

原文摘要 · Abstract (English)

For highly capable AI systems to operate safely in dynamic, open-ended environments, they must be able to identify, understand, and respond to moral reasons for action, and constrain their behaviour accordingly. A growing body of research aims to evaluate this capacity -- moral competence -- in today's most capable AI systems, recently reaching broadly pessimistic conclusions. One of the most ambitious such papers collects gold-standard human-authored rubrics for evaluating moral reasoning in 1,000 cases, and benchmarks frontier AI models against those rubrics, with underwhelming results. In this paper, we argue that the MoReBench dataset can be redeployed to give a much more optimistic picture of LLMs' moral reasoning (an essential part of moral competence). We show that if, instead of scoring LLMs' responses to these cases against these rubrics, we instead give the LLMs the same task given to humans -- to generate scoring rubrics for the moral analysis of particular cases -- the rubrics they generate are both better calibrated to the human rubrics than their open-ended responses, and, where they differ, plausibly reflect nothing more than the vast dimensionality of most moral problems, as well as highlighting some human departures from the "rubric for creating rubrics". Taking these points into consideration, the MoReBench dataset suggests that LLMs are significantly more capable at moral reasoning than was previously believed.

大模型道德推理评估方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。