构建1000个道德困境测试集,评估大模型的推理过程而非仅结果。
MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes
- 设计包含2.3万条评估标准的道德推理测试集,覆盖考量、权衡与建议。
- 发现模型在道德推理上表现不佳,且倾向特定伦理框架。
- 适合关注AI决策透明性与安全性的研究人员使用。
随着人工智能系统能力提升,我们越来越依赖它们协助或替代人类做决策。为确保这些决策符合人类价值观,不仅要关注结果,更要理解其决策过程。具备中间推理痕迹的语言模型为研究这一问题提供了契机。与数学和代码问题不同,道德困境允许多种合理结论,是过程评估的理想场景。为此,我们提出MoReBench:1,000个道德情境,每个配以专家认定的关键评估标准,涵盖识别道德因素、权衡利弊、提出可操作建议等,共包含超过23,000条标准,覆盖人机协同与自主决策等多种场景。同时,我们还构建了MoReBench-Theory:150个案例,用于测试模型在五大学派规范伦理框架下的推理能力。结果显示,传统数学、编程与科学推理任务中的缩放规律无法预测模型在道德推理上的表现;模型对特定伦理框架(如边沁功利主义与康德义务论)存在偏倚,可能是主流训练范式导致的副作用。这些基准推动了以过程为导向的推理评估发展,助力更安全、透明的AI。
原文摘要 · Abstract (English)
As AI systems progress, we rely more on them to make decisions with us and for us. To ensure that such decisions are aligned with human values, it is imperative for us to understand not only what decisions they make but also how they come to those decisions. Reasoning language models, which provide both final responses and (partially transparent) intermediate thinking traces, present a timely opportunity to study AI procedural reasoning. Unlike math and code problems which often have objectively correct answers, moral dilemmas are an excellent testbed for process-focused evaluation because they allow for multiple defensible conclusions. To do so, we present MoReBench: 1,000 moral scenarios, each paired with a set of rubric criteria that experts consider essential to include (or avoid) when reasoning about the scenarios. MoReBench contains over 23 thousand criteria including identifying moral considerations, weighing trade-offs, and giving actionable recommendations to cover cases on AI advising humans moral decisions as well as making moral decisions autonomously. Separately, we curate MoReBench-Theory: 150 examples to test whether AI can reason under five major frameworks in normative ethics. Our results show that scaling laws and existing benchmarks on math, code, and scientific reasoning tasks fail to predict models' abilities to perform moral reasoning. Models also show partiality towards specific moral frameworks (e.g., Benthamite Act Utilitarianism and Kantian Deontology), which might be side effects of popular training paradigms. Together, these benchmarks advance process-focused reasoning evaluation towards safer and more transparent AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。