现有AI道德评估只看对错,忽略了情境判断能力。
Position: Evaluations of AI Moral Reasoning Still Miss Half of the Picture
- 提出道德规范问题,强调模型需理解情境化规则
- 指出当前评估缺失高质量规范数据与推理过程分析
- 适合研究伦理算法、可信AI的学者参考
近期对大语言模型道德能力的评估主要聚焦于‘道德价值问题’——即模型输出是否符合人类道德价值观。相比之下,‘道德规范问题’——模型能否识别并正确应用情境敏感的道德规范——仍被严重忽视。我们指出,这一失衡源于领域对描述性伦理框架(如道德基础理论和科尔伯格道德发展阶段)的依赖,这些框架更关注价值表征而非规范应用。通过回顾现有基准与评估方法,我们发现其高度集中于价值问题,而规范伦理讨论明显不足。我们识别出三个关键缺口:(i) 缺乏高质量的道德规范及其应用的真实标注数据;(ii) 对中间推理过程的评估不足;(iii) 忽视上下文中的道德相关特征识别。为此,我们提出研究议程:包括建立规范理论的标准形式表示、构建专家标注的规范应用数据集,以及明确区分价值层与规范层能力的评估协议。目标是推动对大模型规范推理的系统性研究。
原文摘要 · Abstract (English)
Recent work on evaluating the moral competence of large language models (LLMs) has focused primarily on what we call the moral value problem, i.e., whether model outputs align with human moral values. In contrast, the moral norm problem, i.e., whether models can identify and correctly apply context-sensitive moral norms, remains underexplored. We posit that this imbalance stems from the field's reliance on descriptive ethics frameworks, such as Moral Foundations Theory and Kohlberg's stages of moral development, which emphasize value representation over normative application. We review existing benchmarks and evaluation methods, and show that they cluster heavily around the value problem, while discussion regarding normative ethics remains underrepresented. We identify three crucial gaps: (i) the absence of high-quality ground-truth data for moral norms and their applications, (ii) insufficient evaluation of intermediate reasoning processes, and (iii) limited attention to the identification of morally relevant features in context. Subsequently, we propose a research agenda that includes the development of standardized formal representations for normative theories, the construction of expert-annotated datasets capturing norm application, and evaluation protocols that explicitly distinguish between values-level and norms-level competence. Our goal is to encourage a more systematic study of normative reasoning in LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。