基于费舍尔-罗晓几何改进PPO,实现理论保证的稳定优化。
PPO in the Fisher-Rao geometry
- 用费舍尔-罗晓几何推导更紧的代理目标函数
- 证明策略单调提升,且收敛速度亚线性无维度依赖
- 兼顾理论严谨性与实际表现,适合重视可靠性的强化学习研究者
近端策略优化(PPO)因其出色的实证性能被广泛应用于强化学习,但缺乏策略改进和收敛性的严格保证。PPO的裁剪代理目标源于平坦几何下价值函数线性化的下界。本文通过引入费舍尔-罗晓(Fisher-Rao, FR)几何,推导出更紧的代理目标,并提出费舍尔-罗晓PPO(FR-PPO)。该方法提供强理论保证,包括策略单调改进。在直接参数化设置下,FR-PPO实现亚线性收敛,且不依赖动作或状态空间维度;对于参数化策略,进一步获得亚线性收敛,误差由兼容函数近似误差决定。尽管主要聚焦理论分析,实验也表明FR-PPO在多个标准强化学习任务中表现良好。
原文摘要 · Abstract (English)
Proximal Policy Optimization (PPO) is widely used in reinforcement learning due to its strong empirical performance, yet it lacks formal guarantees for policy improvement and convergence. PPO's clipped surrogate objective is motivated by a lower bound on linearization of the value function in flat geometry setting. We derive a tighter surrogate objective and introduce Fisher-Rao PPO (FR-PPO) by leveraging the Fisher-Rao (FR) geometry. Our scheme provides strong theoretical guarantees, including monotonic policy improvement. In the direct parametrization setting, we show that FR-PPO achieves sub-linear convergence with no dependence on action or state space dimensions, and for parametrized policies we further obtain sub-linear convergence up to the compatible function approximation error. Finally, although our primary focus is theoretical, we also demonstrate empirically that FR-PPO performs well across a range of standard reinforcement learning tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。