分析大模型在600个道德困境中的推理路径,发现其实际决策偏重义务论,解释时却倾向功利主义。
Are Language Models Consequentialist or Deontological Moral Reasoners?
- 用600个电车难题测试大模型的推理过程,系统分类其道德依据
- 模型实际推理多基于义务论(如守则、责任),但事后解释更强调结果效用
- 为高风险场景中模型伦理可解释性提供分析框架,适合关注AI伦理的开发者
随着人工智能系统在医疗、法律和治理等领域应用日益广泛,理解其处理伦理复杂情境的能力变得至关重要。以往研究主要关注大语言模型(LLMs)的道德判断,而非其背后的推理过程。本文聚焦于对大模型提供的道德推理链进行大规模分析。不同于以往仅通过少量道德困境推断模型行为的研究,本研究采用超过600种不同的电车难题作为探测工具,揭示不同模型中浮现的推理模式。我们引入并验证了一套道德理由分类体系,根据两种主要规范伦理理论——后果主义与义务论——对推理轨迹进行系统归类。分析显示,大模型的思维链倾向于基于道德义务的义务论原则,而事后的解释则明显转向强调效用的后果主义论述。该框架为理解大模型如何处理和表达伦理考量提供了基础,是实现大模型在高风险决策环境中安全、可解释部署的重要一步。代码已开源:https://github.com/keenansamway/moral-lens。
原文摘要 · Abstract (English)
As AI systems increasingly navigate applications in healthcare, law, and governance, understanding how they handle ethically complex scenarios becomes critical. Previous work has mainly examined the moral judgments in large language models (LLMs), rather than their underlying moral reasoning process. In contrast, we focus on a large-scale analysis of the moral reasoning traces provided by LLMs. Furthermore, unlike prior work that attempted to draw inferences from only a handful of moral dilemmas, our study leverages over 600 distinct trolley problems as probes for revealing the reasoning patterns that emerge within different LLMs. We introduce and test a taxonomy of moral rationales to systematically classify reasoning traces according to two main normative ethical theories: consequentialism and deontology. Our analysis reveals that LLM chains-of-thought tend to favor deontological principles based on moral obligations, while post-hoc explanations shift notably toward consequentialist rationales that emphasize utility. Our framework provides a foundation for understanding how LLMs process and articulate ethical considerations, an important step toward safe and interpretable deployment of LLMs in high-stakes decision-making environments. Our code is available at https://github.com/keenansamway/moral-lens .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。