arXiv:2409.19437cs.LGcs.AI2024-09被引 4

提出新判据,让策略梯度法在多项式时间内收敛并验证解的优劣。

Strongly-polynomial time and validation analysis of policy gradient methods

  • 用优势差距函数设计步长规则,实现与最优策略分布无关的线性收敛。
  • 在随机设置下,各状态收敛呈次线性,且可近似最优间隙。
  • 提供可计算的最优性验证手段,告别依赖基线对比的盲目评估。

本文为有限状态与动作的马尔可夫决策过程(MDP)和强化学习(RL)提出一种新型终止准则——优势差距函数。通过将该函数融入步长规则设计,并推导出不依赖于最优策略平稳分布的新型线性收敛速率,证明了策略梯度方法可在强多项式时间内求解MDP。据我们所知,这是首次为策略梯度方法建立此类强收敛性质。此外,在仅能获取策略梯度随机估计的随机设定下,优势差距函数能对每个状态的最优性差距提供良好近似,并在每个状态上表现出次线性收敛率。该函数在随机情况下易于估计,结合策略值的易计算上界,可为策略梯度方法生成的解提供便捷的验证方式。因此,我们的研究为强化学习提供了原则性且可计算的最优性度量,而当前实践多依赖算法间或与基线的比较,缺乏最优性保证。

原文摘要 · Abstract (English)

This paper proposes a novel termination criterion, termed the advantage gap function, for finite state and action Markov decision processes (MDP) and reinforcement learning (RL). By incorporating this advantage gap function into the design of step size rules and deriving a new linear rate of convergence that is independent of the stationary state distribution of the optimal policy, we demonstrate that policy gradient methods can solve MDPs in strongly-polynomial time. To the best of our knowledge, this is the first time that such strong convergence properties have been established for policy gradient methods. Moreover, in the stochastic setting, where only stochastic estimates of policy gradients are available, we show that the advantage gap function provides close approximations of the optimality gap for each individual state and exhibits a sublinear rate of convergence at every state. The advantage gap function can be easily estimated in the stochastic case, and when coupled with easily computable upper bounds on policy values, they provide a convenient way to validate the solutions generated by policy gradient methods. Therefore, our developments offer a principled and computable measure of optimality for RL, whereas current practice tends to rely on algorithm-to-algorithm or baselines comparisons with no certificate of optimality.

强化学习策略梯度收敛分析最优性验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。