提出一种新算法,让智能体在复杂环境中更稳定地学习长期平均收益。
Average-reward reinforcement learning in semi-Markov decision processes via relative value iteration
- 基于异步随机逼近思想改进经典算法,适用于弱连通半马尔可夫决策过程。
- 算法几乎必然收敛到最优平均奖励方程的解集,特定条件下收敛到唯一解。
- 新增单调性条件提升稳定性,适合研究强化学习收敛性的学者参考。
本文将作者近期在Borkar-Meyn框架下关于异步随机逼近(SA)的研究成果应用于平均奖励半马尔可夫决策过程(SMDPs)的强化学习。我们建立了异步SA版本的Schweitzer经典相对值迭代算法(RVI Q-learning)在有限状态空间、弱连通SMDPs下的收敛性。特别地,证明了该算法几乎必然收敛至平均奖励最优方程的紧致连通解集,在额外步长与异步性条件下可进一步收敛至唯一、路径依赖的解。此外,为充分运用SA框架,我们引入了估计最优奖励率的新单调性条件,显著扩展了先前算法框架,并通过新颖的稳定性与收敛性分析加以论证。
原文摘要 · Abstract (English)
This paper applies the authors' recent results on asynchronous stochastic approximation (SA) in the Borkar-Meyn framework to reinforcement learning in average-reward semi-Markov decision processes (SMDPs). We establish the convergence of an asynchronous SA analogue of Schweitzer's classical relative value iteration algorithm, RVI Q-learning, for finite-space, weakly communicating SMDPs. In particular, we show that the algorithm converges almost surely to a compact, connected subset of solutions to the average-reward optimality equation, with convergence to a unique, sample path-dependent solution under additional stepsize and asynchrony conditions. Moreover, to make full use of the SA framework, we introduce new monotonicity conditions for estimating the optimal reward rate in RVI Q-learning. These conditions substantially expand the previously considered algorithmic framework and are addressed through novel arguments in the stability and convergence analysis of RVI Q-learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。