用残差加权修正加速Q-learning,理论证明收敛更快。
Heavy-Ball Q-Learning with Residual Weighting Correction

- 引入残差加权的动量机制改进Q-learning
- 理论证明在特定条件下收敛速度优于标准Q-learning
- 适用于需要快速收敛的强化学习场景
本文提出一种修正的Heavy-Ball Q-learning方法,并建立了其确定性均值动态的收敛性。同时识别出该方法在理论上可比标准Q-learning更快收敛的条件。该构造进一步拓展至线性函数逼近的Q-learning,推导出对应修正不动点的收敛与加速性质。通过条件均值递推分析采样随机版本,在所述线性函数逼近设定下获得有限时间界。分析基于Q-learning算法的切换线性系统(SLS)表示及其关联切换族的联合谱半径(JSR)。这一SLS视角在标准Q-learning分析中并不常见,为理解动量如何加速Q-learning提供了互补框架与新见解。
原文摘要 · Abstract (English)
This paper proposes a corrected heavy-ball Q-learning method for reinforcement learning (RL) and establishes convergence of its deterministic mean dynamics. It also identifies conditions under which the method is theoretically guaranteed to converge faster than standard Q-learning. The same construction is then extended to Q-learning with linear function approximation, where analogous convergence and acceleration statements are derived for the corresponding corrected fixed point. The sampled stochastic versions are treated through conditional-mean recursions and, in the stated linear-function-approximation setting, finite-time bounds. The analysis is based on a switched linear system (SLS) representation of Q-learning algorithms and on the joint spectral radius (JSR) of the associated switching families. This SLS viewpoint is not commonly used in standard analyses of Q-learning, and it provides a complementary framework and new insight into how heavy-ball momentum can accelerate Q-learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。