arXiv:2608.19587cs.LG2026-08

提出单循环熵正则化自然策略梯度算法,实现无正则化目标的加速收敛。

Unregularized Convergence of Single-Loop, Entropy-Regularized Natural Actor-Critic

  • 采用单循环结构与未中心化评论家,稳定应对策略趋确定性时的奇异信息矩阵。
  • 在随机与确定性场景下分别达到 $\tilde{\mathcal{O}}(T^{-1})$ 与 $\tilde{\mathcal{O}}(T^{-2/3})$ 收敛率。
  • 适用于带正动作间隙的MDP,突破传统统计下界,适合强化学习理论研究者。

尽管熵正则化广泛用于稳定和加速自然策略梯度方法,但其对无正则化目标的加速收敛能力仍缺乏深入探索。现有分析多依赖双循环架构并假设线性熵惩罚。为弥合理论与实践差距,本文在兼容线性函数近似下分析一种单循环、熵正则化的自然演员-评论家算法。通过训练未中心化评论家,其跟踪过程在策略趋确定性及费舍尔信息矩阵退化时仍保持稳定。针对优化景观的两类主要情形:随机情形下将耦合的演员-评论家更新融合为联合李雅普诺夫递推;确定性情形下转向策略镜面下降框架以规避欧几里得几何坍塌。利用无正则化马尔可夫决策过程中的正最小动作间隙,引入指数平移机制,将正则化间隙映射至无正则化间隙,尾部呈指数衰减。通过固定温度调节,算法实现加速的无正则化收敛率:随机情形下为 $\tilde{\mathcal{O}}(T_{total}^{-1})$,确定性情形下平均迭代器达 $\tilde{\mathcal{O}}(T_{total}^{-2/3})$,最后迭代器达 $\tilde{\mathcal{O}}(T_{total}^{-1/3})$,其中 $T_{total}$ 为总随机评论家更新次数(或蒙特卡洛回溯次数)。此外,在表格设置下,正动作间隙分析给出平均迭代器 $\tilde{\mathcal{O}}(T_{total}^{-2/3})$ 的速率,超越无正动作裕度时适用的 $\mathcal{O}(T_{total}^{-1/2})$ 最坏情况统计下界。

原文摘要 · Abstract (English)

While entropy regularization is widely used to stabilize and accelerate Natural Policy Gradient methods, its ability to yield faster convergence rates for the unregularized objective remains underexplored. Existing analyses often rely on double-loop architectures and invoke a linear entropy penalty. To bridge the gap between theory and practice, we analyze a single-loop, entropy-regularized Natural Actor-Critic algorithm under compatible linear function approximation. By training an uncentered critic, our critic tracking can remain stable even as the training policy approaches determinism and the Fisher information matrix degenerates. We focus on two primary regimes for the optimization landscape: a Stochastic Regime, where we fuse coupled actor-critic updates into a joint Lyapunov recurrence, and a Deterministic Regime, where we pivot to a Policy Mirror Descent framework to circumvent the collapse of Euclidean geometry. By exploiting a positive Minimal Action Gap in the unregularized Markov decision process, we introduce an Exponential Translation mechanism that maps the regularized gap to the unregularized one up to an exponentially decaying tail. By tuning the fixed temperature, our algorithm achieves accelerated unregularized convergence rates, up to approximation-error terms: $\tilde{\mathcal{O}}(T_{total}^{-1})$ in the Stochastic Regime, and $\tilde{\mathcal{O}}(T_{total}^{-2/3})$ for the average iterate alongside $\tilde{\mathcal{O}}(T_{total}^{-1/3})$ for the last iterate in the Deterministic Regime. Here, $T_{total}$ denotes the total number of stochastic critic updates (or Monte Carlo rollouts). Furthermore, in the tabular setting, our positive-action-gap analysis yields a $\tilde{\mathcal{O}}(T_{total}^{-2/3})$ average-iterate rate, surpassing the $\mathcal{O}(T_{total}^{-1/2})$ worst-case statistical barrier that applies without a positive action margin.

强化学习策略优化收敛分析熵正则化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。