arXiv:2605.15960cs.AIcs.LG2026-05被引 1

世界模型存在漏洞时,会导致智能体误判策略优劣,难避免被利用。

Imperfect World Models are Exploitable

论文配图:Imperfect World Models are Exploitable
图 1 · 摘自论文原文
  • 提出模型利用的新定义:模型推荐的策略实际比环境真实情况差。
  • 证明在大规模策略集中,模型利用几乎不可避免,且奖励黑客行为是其特例。
  • 引入宽松定义并给出安全规划时间范围,为实际应用提供边界参考。

我们提出了强化学习中模型利用的新定义。直观而言,若一个世界模型暗示某一策略应严格优于另一策略,而环境的真实转移模型却显示相反,则该模型即为可被利用的。我们将此定义类比于奖励黑客的先前刻画,但指出其必然性证明无法直接推广至模型利用。为此,我们构建了奖励黑客与模型利用的一般理论,证明在大规模策略集上,模型利用本质上不可避免,并推出黑客行为作为特例同样无法避免。然而,我们还发现,有限策略集下保证无黑客的条件,无法对应出防止利用的条件。因此,我们引入一种放松的利用概念,并推导出可避免利用的安全规划时间范围。综合来看,我们的结果在奖励黑客与模型利用之间建立了正式桥梁,阐明了世界模型安全规划的局限性。

原文摘要 · Abstract (English)

We propose a novel definition of model exploitation in reinforcement learning. Informally, a world model is exploitable if it implies that one policy should be strictly preferred over another while the environment's true transition model implies the reverse. We analogize our definition with a prior characterization of reward hacking but show that the associated proof of inevitability does not transfer to exploitation. To overcome this obstruction, we develop a general theory of reward hacking and model exploitation that proves that exploitation is essentially unavoidable on large policy sets and yields the corresponding claim for hacking as a special case. Unfortunately, we also find that the conditions that guarantee unhackability in finite policy sets have no counterpart that precludes exploitation. Consequently, we introduce a relaxed notion of exploitation and derive a safe horizon within which it can be avoided. Taken together, our results establish a formal bridge between reward hacking and model exploitation and elucidate the limits of safe planning in world models.

强化学习模型利用安全规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。