证明了Wasserstein策略优化在连续空间中线性收敛。
A note on convergence of Wasserstein policy optimization
- 基于熵正则马尔可夫决策过程,利用对数索博列夫不等式分析梯度流。
- 证明策略优化过程中能量单调衰减,值函数线性逼近全局最优。
- 适合研究强化学习理论收敛性的学者参考。
Wasserstein策略优化(WPO)是一种最近提出的强化学习算法,利用Wasserstein梯度流在连续动作空间中优化随机策略。尽管其在实践中表现良好,但其在连续状态与动作空间环境中的理论收敛性尚未完全建立。本文指出,在熵正则马尔可夫决策过程框架下,WPO具有线性收敛性。通过借助近期关于梯度流收敛的均场分析及对数索博列夫不等式,假设梯度流方程存在足够光滑的解,我们证明了沿流的能量单调衰减,并建立了局部对数索博列夫不等式。这些性质最终使我们能够论证值函数将线性收敛至全局最优。
原文摘要 · Abstract (English)
Wasserstein Policy Optimization (WPO) is a recently proposed reinforcement learning algorithm that leverages Wasserstein gradient flows to optimize stochastic policies in continuous action spaces. Despite its empirical success, the theoretical convergence properties of WPO in environments with continuous state and action spaces have yet to be fully established. In this note, we argue that WPO within the framework of entropy-regularised Markov Decision Processes converges linearly. This is done by leveraging recent advances in mean-field analysis for convergence of gradient flows using log-Sobole inequalities. Assuming existence of sufficiently regular solution to the gradient flow equation we demonstrate monotonic energy dissipation along the flow and establish a local log-Sobolev inequality. Ultimately, these properties allow us to argue that the value function should converge linearly to the global optimum.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。