用世界模型实现安全策略改进,兼顾理论保证与在线学习性能。
Deep SPI: Safe Policy Improvement via World Models
- 基于世界模型构建局部策略更新机制,确保学习过程稳定
- 在ALE-57上表现媲美甚至超越PPO和DeepMDPs等强基线
- 首次将离线强化学习的理论保障扩展到在线深度强化学习
安全策略改进(SPI)提供了对策略更新的理论控制,但现有保证主要局限于离线、表格型强化学习。本文研究在结合世界模型与表征学习的通用在线设置下进行SPI。我们建立了一个理论框架,表明限制策略更新在当前策略的明确定义邻域内,可保证单调改进与收敛性。该分析将转移与奖励预测损失与表征质量关联,得到离线强化学习中经典SPI定理的在线、深度版本。基于此,我们提出DeepSPI,一种基于局部转移与奖励损失及正则化策略更新的系统性在线算法。在ALE-57基准上,DeepSPI达到或超过强基线(包括PPO和DeepMDPs),同时保持理论保证。
原文摘要 · Abstract (English)
Safe policy improvement (SPI) offers theoretical control over policy updates, yet existing guarantees largely concern offline, tabular reinforcement learning (RL). We study SPI in general online settings, when combined with world model and representation learning. We develop a theoretical framework showing that restricting policy updates to a well-defined neighborhood of the current policy ensures monotonic improvement and convergence. This analysis links transition and reward prediction losses to representation quality, yielding online, "deep" analogues of classical SPI theorems from the offline RL literature. Building on these results, we introduce DeepSPI, a principled on-policy algorithm that couples local transition and reward losses with regularised policy updates. On the ALE-57 benchmark, DeepSPI matches or exceeds strong baselines, including PPO and DeepMDPs, while retaining theoretical guarantees.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。