arXiv:2605.24939cs.LGmath.OC2026-05

证明了连续空间下带熵正则的策略梯度全局线性收敛。

Global linear convergence of entropy-regularized softmax policy gradient beyond tabular MDPs

  • 采用线性函数逼近的软最大策略,扩展了表格型设置。
  • 在特定特征条件下,策略梯度使误差以指数速度下降。
  • 适用于强化学习中复杂连续环境的理论分析。

我们研究无限时域熵正则化马尔可夫决策过程(MDPs)在连续状态与动作空间下的策略梯度全局收敛性。考虑使用线性函数逼近的对数线性软最大策略,该参数化在保留可处理性的同时扩展了表格型软最大参数化。在正则化状态-动作值函数 $Q^π_τ$-可实现的前提下,我们首先建立了一个非均匀的 Polyak--Łojasiewicz(PŁ)不等式,其非均匀性源于与策略几何相关的常数退化,即费舍尔信息矩阵或未中心化特征协方差矩阵。随后,我们识别出两种特征情形,在这些情形下该非均匀常数可在梯度流路径上被有界。对于全仿射张成特征,我们证明了KL正则项的径向无界性,并表明费舍尔信息矩阵的最小特征值始终不低于一个依赖于初始化的正常数。对于单纯形取值特征,我们在与全1向量正交的子空间中证明了类似的径向无界性,并获得了未中心化协方差矩阵最小特征值的统一下界。这些结果表明,正则化目标沿梯度流实现全局线性收敛,即次优性以 $\\(\mathcal{O}(e^{-Ct})$ 的形式衰减,其中 $C>0$。我们的分析将 Agarwal 等 (2020);Bhandari 与 Russo (2024);Mei 等 (2020) 的表格型设定下的全局收敛理论推广至连续空间。

原文摘要 · Abstract (English)

We study the global convergence of policy gradient for infinite-horizon entropy-regularized Markov decision processes (MDPs) with continuous state and action spaces. We consider log-linear softmax policies with linear function approximation, which extend the tabular softmax parameterization while retaining a tractable policy class. Under $Q^π_τ$-realizability for the regularized state-action value function, we first establish a non-uniform Polyak--Łojasiewicz (PŁ) inequality. The non-uniformity arises through degeneracy of constants associated with the policy geometry, namely the Fisher information matrix or an uncentered feature covariance matrix. We then identify two feature regimes under which this non-uniform constant can be bounded along the gradient flow. For full-affine-span features, we prove radial unboundedness of the KL regularizer and show that the smallest eigenvalue of the Fisher information matrix remains bounded below by an initialization-dependent positive constant. For simplex-valued features, we prove an analogous radial unboundedness result in the subspace orthogonal to the all-ones vector and obtain a uniform lower bound for the smallest eigenvalue of the uncentered covariance matrix. These results imply global linear convergence of the regularized objective along the gradient flow, i.e. suboptimality decaying as $\mathcal{O}(e^{-Ct})$ for some $C>0$. Our analysis extends the global convergence theory of entropy-regularized softmax policy gradient beyond the tabular setting of Agarwal et al. (2020); Bhandari and Russo (2024); Mei et al. (2020).

强化学习策略梯度收敛性分析熵正则

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。