构建首个评估智能体道德决策的基准测试,聚焦复杂情境下的伦理权衡。
MoralityGym: A Benchmark for Evaluating Hierarchical Moral Alignment in Sequential Decision-Making Agents
- 提出道德链条形式化框架,将规范按层级约束建模
- 设计98个类电车难题环境,分离任务执行与道德评估
- 引入新度量标准,融合哲学心理学洞察评估道德敏感性
在人工智能安全、伦理哲学与认知科学交汇处,评估智能体在冲突且分层的人类规范中做出道德决策的能力是一项关键挑战。本文提出Morality Chains——一种将道德规范表示为有序义务约束的新形式化方法,并构建了包含98个伦理困境的MoralityGym基准,以类电车难题风格的Gymnasium环境呈现。通过将任务求解与道德评估解耦,并引入新的道德度量标准,该基准支持整合心理学与哲学洞见,评估智能体对规范敏感的推理能力。基于安全强化学习的基线结果揭示了现有方法的关键局限,凸显出发展更严谨伦理决策方法的必要性。本研究为开发在复杂现实场景中更可靠、透明且合乎伦理的AI系统奠定了基础。
原文摘要 · Abstract (English)
Evaluating moral alignment in agents navigating conflicting, hierarchically structured human norms is a critical challenge at the intersection of AI safety, moral philosophy, and cognitive science. We introduce Morality Chains, a novel formalism for representing moral norms as ordered deontic constraints, and MoralityGym, a benchmark of 98 ethical-dilemma problems presented as trolley-dilemma-style Gymnasium environments. By decoupling task-solving from moral evaluation and introducing a novel Morality Metric, MoralityGym allows the integration of insights from psychology and philosophy into the evaluation of norm-sensitive reasoning. Baseline results with Safe RL methods reveal key limitations, underscoring the need for more principled approaches to ethical decision-making. This work provides a foundation for developing AI systems that behave more reliably, transparently, and ethically in complex real-world contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。