证明了熵正则强化学习中Wasserstein策略梯度的全局收敛性。
Global Convergence of Wasserstein Policy Gradient for Entropy-Regularized Reinforcement Learning
- 利用贝尔曼结构构建非凸问题的PL型几何,替代传统凸性假设。
- 在均匀对数索博列夫不等式下实现指数级收敛,误差受离散化偏差限制。
- 适用于连续控制任务,为基于最优传输的策略优化提供理论支持。
Wasserstein策略梯度(WPG)是一种利用动作分布最优传输几何的强化学习策略优化方法。针对熵正则化强化学习目标,WPG通过沿软Q函数的动作梯度并结合朗之万型扩散来演化每个状态条件下的策略。尽管该方法在连续控制问题中表现优异,其全局收敛性仍缺乏理论理解。标准朗之万分析不适用,因为强化学习目标通过贝尔曼递归依赖于策略,而非静态凸泛函,且朗之万漂移由软Q函数决定,其正则性需在策略迭代过程中控制。本文通过利用熵正则化强化学习的贝尔曼结构,建立了一套全局收敛理论。我们发现,通常由凸性扮演的角色可被贝尔曼基论证替代:软贝尔曼残差相对于吉布斯策略具有状态层面的KL表示;贝尔曼收缩将该残差与全局最优差距关联;贝尔曼求解恒等式将价值提升与相对费舍尔信息联系起来。结合演化吉布斯族的均匀对数索博列夫不等式(LSI),这些要素共同导出分布式的Polyak–Łojasiewicz条件。进一步建立了控制离散化误差所需的正则性和统一有界性,从而获得几何收缩至离散化偏差。概念上,我们的分析表明,虽然熵正则化强化学习在传统平坦意义下并非凸,但贝尔曼递归诱导出有利的Polyak–Łojasiewicz型(PL)几何,支持WPG的全局收敛。
原文摘要 · Abstract (English)
Wasserstein policy gradient (WPG) is a policy optimization method for reinforcement learning (RL) that exploits the optimal-transport geometry of action distributions. For the entropy-regularized RL objective, WPG evolves each state-conditional policy by transporting it along the action gradient of the soft Q-function together with a Langevin-type diffusion. Despite its appeal for continuous-control problems, its global convergence properties remain poorly understood. Standard Langevin analyses do not directly apply, because the RL objective depends on the policy through the Bellman recursion rather than through a static convex functional, and the Langevin drift is determined by the soft Q-function, whose regularity must be controlled along the policy iterates. In this paper, we develop a global convergence theory for WPG by exploiting the Bellman structure of entropy-regularized RL. We show that the role usually played by convexity can be replaced by a Bellman-based argument: the soft Bellman residual admits a statewise KL representation with respect to a Gibbs policy; Bellman contraction relates this residual to the global optimality gap; and a Bellman resolvent identity connects value improvement to relative Fisher information. Combined with a uniform log-Sobolev inequality (LSI) for the evolving Gibbs family, these ingredients yield a distributional Polyak--Łojasiewicz condition. We further establish the regularity and uniform bounds needed to control the discretization error, thereby obtaining geometric contraction up to a discretization bias. Conceptually, our analysis shows that although entropy-regularized RL is not convex in the usual flat sense, the Bellman recursion induces a favorable Polyak--Lojasiewicz-type (PL) geometry that supports global convergence of WPG.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。