提出新型异步随机逼近理论,支撑平均奖励强化学习算法收敛
Asynchronous Stochastic Approximation with Applications to Average-Reward Reinforcement Learning
- 拓展噪声条件下的稳定性证明,支持更广泛的异步算法
- 通过动力系统方法分析算法轨迹,获得更强收敛保证
- 为基于相对值迭代的平均奖励强化学习提供理论基础
本文研究异步随机逼近(SA)算法的稳定性和收敛性,重点拓展其在平均奖励强化学习中的应用。首先将Borkar与Meyn的稳定性证明方法推广至更一般的噪声条件下,从而获得更广泛的收敛保障。为进一步精确收敛分析,借鉴Hirsch和Benaïm的动力系统方法,研究了异步SA的影子性质。这些结果为配套论文中提出的基于相对值迭代的强化学习算法提供了理论基础,用于求解平均奖励马尔可夫决策过程和半马尔可夫决策过程。
原文摘要 · Abstract (English)
This paper investigates the stability and convergence properties of asynchronous stochastic approximation (SA) algorithms, with a focus on extensions relevant to average-reward reinforcement learning. We first extend a stability proof method of Borkar and Meyn to accommodate more general noise conditions than previously considered, thereby yielding broader convergence guarantees for asynchronous SA. To sharpen the convergence analysis, we further examine the shadowing properties of asynchronous SA, building on a dynamical systems approach of Hirsch and Benaïm. These results provide a theoretical foundation for a class of relative value iteration-based reinforcement learning algorithms -- developed and analyzed in a companion paper -- for solving average-reward Markov and semi-Markov decision processes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。