arXiv:2506.13082cs.AI2025-06被引 13

评测大模型的道德能力,发现其在复杂情境下识别关键信息的能力远不如人类。

Discerning What Matters: A Multi-Dimensional Assessment of Moral Competence in LLMs

  • 设计五维评估框架,考察模型识别、权衡、推理和补全道德判断的能力
  • 在普通情境中模型表现优于普通人,但在含干扰信息的新场景中大幅落后
  • 提醒研究者警惕当前评测方法的局限性,适合关注AI伦理的开发者参考

道德能力是依据道德原则行事的能力。随着大语言模型(LLMs)在需道德判断的场景中日益广泛应用,对其能力进行实证评估日益重要。我们回顾现有文献,指出三大缺陷:(i) 过度依赖预设道德情境且明确标注道德要素;(ii) 侧重结论预测而非道德推理过程;(iii) 未能充分测试模型识别信息缺失的能力。基于哲学中的道德技能研究,我们提出一种评估LLM道德能力的新方法。该方法超越简单结论比对,从五个维度评估:识别道德相关特征、权衡其重要性、为特征分配道德理由、整合一致的道德判断、识别信息缺口。我们通过两项实验,比较六种主流LLM与非专家人类及专业哲学家的表现。第一项实验使用标准伦理短篇,LLMs在多个维度上普遍优于非专家人类。但第二项实验采用新设计的情境,通过将相关特征隐藏于无关细节中测试道德敏感性,结果出现显著逆转:部分LLMs表现显著劣于人类。这表明当前评估可能严重高估了LLMs的道德推理能力,因忽略了从噪声中辨识道德相关性的基本要求,而这是真正道德技能的前提。本研究提供更精细的评估框架,并指明提升高级AI道德能力的重要方向。

原文摘要 · Abstract (English)

Moral competence is the ability to act in accordance with moral principles. As large language models (LLMs) are increasingly deployed in situations demanding moral competence, there is increasing interest in evaluating this ability empirically. We review existing literature and identify three significant shortcoming: (i) Over-reliance on prepackaged moral scenarios with explicitly highlighted moral features; (ii) Focus on verdict prediction rather than moral reasoning; and (iii) Inadequate testing of models' (in)ability to recognize when additional information is needed. Grounded in philosophical research on moral skill, we then introduce a novel method for assessing moral competence in LLMs. Our approach moves beyond simple verdict comparisons to evaluate five dimensions of moral competence: identifying morally relevant features, weighting their importance, assigning moral reasons to these features, synthesizing coherent moral judgments, and recognizing information gaps. We conduct two experiments comparing six leading LLMs against non-expert humans and professional philosophers. In our first experiment using ethical vignettes standard to existing work, LLMs generally outperformed non-expert humans across multiple dimensions of moral reasoning. However, our second experiment, featuring novel scenarios designed to test moral sensitivity by embedding relevant features among irrelevant details, revealed a striking reversal: several LLMs performed significantly worse than humans. Our findings suggest that current evaluations may substantially overestimate LLMs' moral reasoning capabilities by eliminating the task of discerning moral relevance from noisy information, which we take to be a prerequisite for genuine moral skill. This work provides a more nuanced framework for assessing AI moral competence and highlights important directions for improving moral competence in advanced AI systems.

道德推理评估框架大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。