arXiv:2511.12271cs.AI2025-11AAAI被引 5

让大模型学会在新情境中坚持特定道德准则,提升决策可靠性。

MoralReason: Generalizable Moral Decision Alignment For LLM Agents Using Reasoning-Level Reinforcement Learning

  • 用推理层级强化学习训练模型,同时优化决策与道德推理过程。
  • 在未见过的情境中,功利主义和义务论对齐分数分别提升0.757和0.450。
  • 适用于需可解释道德决策的AI系统,如医疗、司法辅助场景。

大型语言模型正日益影响人类道德判断,但现有方法多聚焦评估而非主动引导其道德决策。本文将此问题定义为分布外道德对齐难题,即大模型代理需在训练数据之外的新情境中应用一致的道德推理框架。为此,我们构建了Moral-Reason-QA数据集,扩展了680个由人类标注的高模糊性道德场景,并加入功利主义、义务论与美德伦理三类框架的推理轨迹,实现对真实决策情境下道德泛化能力的系统评估。所提学习方法采用分组相对策略优化,结合复合奖励函数,同步优化决策对齐与框架特异性推理过程,以促进内在道德框架的学习。实验表明,模型在分布外测试集上成功泛化,软最大归一化对齐分数在功利主义框架下提升0.757,在义务论框架下提升0.450。实验还揭示了训练挑战与未来研究方向。结果证明,大模型可被系统训练以内化并应用于新情境的特定道德框架,为语言模型深度融入人类决策过程提供关键安全基础。

原文摘要 · Abstract (English)

Large language models are increasingly influencing human moral decisions, yet current approaches focus primarily on evaluating rather than actively steering their moral decisions. We formulate this as an out-of-distribution moral alignment problem, where LLM agents must learn to apply consistent moral reasoning frameworks to scenarios beyond their training distribution. We introduce Moral-Reason-QA, a novel dataset extending 680 human-annotated, high-ambiguity moral scenarios with framework-specific reasoning traces across utilitarian, deontological, and virtue ethics, enabling systematic evaluation of moral generalization in realistic decision contexts. Our learning approach employs Group Relative Policy Optimization with composite rewards that simultaneously optimize decision alignment and framework-specific reasoning processes to facilitate learning of the underlying moral frameworks. Experimental results demonstrate successful generalization to unseen moral scenarios, with softmax-normalized alignment scores improving by +0.757 for utilitarian and +0.450 for deontological frameworks when tested on out-of-distribution evaluation sets. The experiments also reveal training challenges and promising directions that inform future research. These findings establish that LLM agents can be systematically trained to internalize and apply specific moral frameworks to novel situations, providing a critical foundation for AI safety as language models become more integrated into human decision-making processes.

道德对齐强化学习大模型推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。