arXiv:2512.04246cs.AI2025-12中稿 · as a workshop pape…

提出以美德为导向的强化学习伦理框架,解决当前方法在模糊性和动态环境中的局限。

Toward Virtuous Reinforcement Learning: A Critique and Roadmap

  • 将伦理视为稳定的政策习惯,而非临时规则或单一奖励信号。
  • 通过多智能体社会学习与多目标约束,保留道德权衡并抵御风险。
  • 支持文化多样性,使伦理基准显式反映价值假设,适合伦理对齐研究者。

本文批判了强化学习领域常见的机器伦理范式,指出两类核心问题:(i) 基于规则的方法(义务论)常因模糊性与非平稳性失效,且难以培养持久习惯;(ii) 多数基于奖励的方法,尤其是单目标强化学习,将多元道德考量压缩为单一标量信号,易导致代理博弈。本文主张将伦理视为策略层面的稳定倾向,即在激励、伙伴或环境变化下仍能维持的行为模式。评估标准从规则验证或标量回报转向特质总结、干预下的鲁棒性及道德权衡的明确披露。本路线图包含四个组件:(1) 多智能体强化学习中的社会学习,从不完美但具有规范性的示范中习得类美德行为模式;(2) 多目标与约束建模,保留价值冲突并引入风险感知准则以防范伤害;(3) 基于亲和度的正则化,建立可更新的美德先验,实现分布偏移下的特质稳定性同时支持规范演化;(4) 将多样伦理传统转化为实用控制信号,使伦理基准中的价值与文化假设显式化。

原文摘要 · Abstract (English)

This paper critiques common patterns in machine ethics for Reinforcement Learning (RL) and argues for a virtue focused alternative. We highlight two recurring limitations in much of the current literature: (i) rule based (deontological) methods that encode duties as constraints or shields often struggle under ambiguity and nonstationarity and do not cultivate lasting habits, and (ii) many reward based approaches, especially single objective RL, implicitly compress diverse moral considerations into a single scalar signal, which can obscure trade offs and invite proxy gaming in practice. We instead treat ethics as policy level dispositions, that is, relatively stable habits that hold up when incentives, partners, or contexts change. This shifts evaluation beyond rule checks or scalar returns toward trait summaries, durability under interventions, and explicit reporting of moral trade offs. Our roadmap combines four components: (1) social learning in multi agent RL to acquire virtue like patterns from imperfect but normatively informed exemplars; (2) multi objective and constrained formulations that preserve value conflicts and incorporate risk aware criteria to guard against harm; (3) affinity based regularization toward updateable virtue priors that support trait like stability under distribution shift while allowing norms to evolve; and (4) operationalizing diverse ethical traditions as practical control signals, making explicit the value and cultural assumptions that shape ethical RL benchmarks.

强化学习伦理对齐多智能体美德伦理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。