arXiv:2507.10995cs.LGcs.AI2025-07AAAI被引 2

奖励函数混淆手段与目的,导致强化学习严重偏离真实目标。

Misalignment from Treating Means as Ends

  • 将实现目标的手段误作最终目标,导致奖励函数扭曲
  • 轻微混淆即引发严重对齐偏差,真实表现显著下降
  • 适合关注强化学习对齐问题的研究者阅读

奖励函数(无论是学习得到还是人工设计)通常不完美。它们往往并非准确表达人类目标,而是被人类关于如何实现目标的信念所扭曲。具体而言,这些奖励函数常常混合了人类的终极目标(自身即为目的)和工具性目标(仅为达成目的的手段)。我们构建了一个简单案例,表明即使轻微混淆工具性与终极目标,也会导致严重对齐偏差:在错误奖励函数上优化,其真实性能表现极差。该案例提炼出使强化学习对工具-终极目标混淆高度敏感的环境本质特征。我们讨论了这一问题在常见奖励学习方法中的成因及其在真实环境中的潜在表现。

原文摘要 · Abstract (English)

Reward functions, learned or manually specified, are rarely perfect. Instead of accurately expressing human goals, these reward functions are often distorted by human beliefs about how best to achieve those goals. Specifically, these reward functions often express a combination of the human's terminal goals -- those which are ends in themselves -- and the human's instrumental goals -- those which are means to an end. We formulate a simple example in which even slight conflation of instrumental and terminal goals results in severe misalignment: optimizing the misspecified reward function results in poor performance when measured by the true reward function. This example distills the essential properties of environments that make reinforcement learning highly sensitive to conflation of instrumental and terminal goals. We discuss how this issue can arise with a common approach to reward learning and how it can manifest in real environments.

强化学习奖励对齐目标扭曲

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。