arXiv:2608.08220cs.AI2026-08

用元规范理论重新审视强化学习道德代理的合理性

Metanormative Theory for RL-Based Moral Agents

  • 从元规范理论中提取设计道德智能体的思路
  • 为强化学习行为提供是否算道德的标准
  • 适合研究机器伦理与价值对齐的学者

机器伦理与价值对齐领域致力于设计符合人类价值观且行为合乎伦理的人工智能代理。近期趋势是使用强化学习(RL)来构建此类代理,而逐渐忽视了以往起核心作用的哲学文献。本文旨在实现两个目标:一是提炼近期元规范理论中的思想,以辅助设计人工道德与价值对齐代理;二是从这些思想视角审视强化学习架构。这将帮助我们明确界定何时可将强化学习代理的行为视为道德行为,并为评估和比较不同的基于强化学习的机器伦理与价值对齐方法提供依据。

原文摘要 · Abstract (English)

The overlapping disciplines of machine ethics and value alignment are concerned with designing artificial agents that are aligned with human values and that act in ethically acceptable ways. A recent trend in these disciplines is the use of reinforcement learning (RL) to design such agents, sidelining the philosophical literature that used to play a more central role. Against this backdrop, this paper pursues two goals. The first is to draw out ideas from recent work in metanormative theory that can be useful for designing artificial moral and value-aligned agents. The second is to examine the RL architecture through the lens of these ideas. This will give us clearer criteria for when an RL agent's behavior can be classified as moral, as well as a basis for evaluating and comparing different RL-based approaches to machine ethics and value alignment.

机器伦理强化学习价值对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。