研究离线多步TD学习在函数逼近下的收敛性,证明其在采样步数足够时能稳定求解。
Analysis of Off-Policy $n$-Step TD-Learning with Linear Function Approximation
- 通过分析确定性模型算法,建立理论基础
- 证明当步数n足够大时算法收敛到有意义解
- 为离线强化学习提供理论支持,适合算法研究者
本文分析了在“致命三重奏”场景下(线性函数逼近、离线学习、自举)的多步时序差分(TD)学习算法。首先,全面研究其基于模型的确定性对应算法,包括投影值迭代和梯度下降算法,这些可视为原型确定性算法,其分析对理解和发展无模型强化学习算法至关重要。特别地,证明当采样步数n足够大时,这些算法收敛至有意义的解。基于上述结论,在第二部分提出了两种多步TD学习算法并进行分析,可视为前述模型基确定性算法的无模型强化学习对应方法。
原文摘要 · Abstract (English)
This paper analyzes multi-step temporal difference (TD)-learning algorithms within the ``deadly triad'' scenario, characterized by linear function approximation, off-policy learning, and bootstrapping. In particular, we prove that $n$-step TD-learning algorithms converge to a solution as the sampling horizon $n$ increases sufficiently. The paper is divided into two parts. In the first part, we comprehensively examine the fundamental properties of their model-based deterministic counterparts, including projected value iteration, gradient descent algorithms, which can be viewed as prototype deterministic algorithms whose analysis plays a pivotal role in understanding and developing their model-free reinforcement learning counterparts. In particular, we prove that these algorithms converge to meaningful solutions when $n$ is sufficiently large. Based on these findings, in the second part, two $n$-step TD-learning algorithms are proposed and analyzed, which can be seen as the model-free reinforcement learning counterparts of the model-based deterministic algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。