解析Q-learning中正负误差的不对称收敛,揭示过估计机制
Sign-Separated Asymmetric Finite-Time Error Analysis of Q-Learning
- 分解Q值误差为正负两部分,分别分析其收敛速率
- 正误差收敛更慢,负误差收敛更快,体现过估计放大效应
- 理论结合实例,为最大值诱导过估计提供直接证据
Q-learning存在过估计偏差:由于Bellman更新对噪声或不准确的动作价值估计进行最大化,正误差可能被选中并传播,导致学习值超过真实最优值。这会减缓学习、降低策略质量并使价值估计不可靠。尽管Q-learning的收敛性已有广泛研究,但能明确反映该过估计机制的收敛理论仍有限。本文研究由过估计偏差引发的Q-learning非对称收敛行为。将Q-learning误差分解为其分量上的正负部分,推导出两者的独立有限时间收敛速率。结果表明,正误差部分可被赋予比负误差更慢的指数衰减速率。这种速率分离为最大值引起的过估计提供了间接理论支持:正误差在最大化步骤中被放大,而负误差则能与最优策略系统进行更严格的比较。该分离基于上界差异,不保证每条轨迹均成立。但我们构造了实际轨迹中呈现预测不对称性的例子。分析给出了确定性和随机常步长下的有限时间界,并阐明了过估计如何影响Q-learning的切换系统动态。
原文摘要 · Abstract (English)
Q-learning is known to suffer from overestimation bias: because the Bellman update maximizes noisy or imperfect action-value estimates, positive errors can be selected and propagated, causing learned values to exceed the true optimal values. This bias can slow learning, degrade policy quality, and make value estimates unreliable. Although the convergence of Q-learning has been studied extensively, convergence theory that explicitly reflects this overestimation mechanism remains limited. This paper studies the asymmetric convergence behavior of Q-learning induced by overestimation bias. We decompose the Q-learning error into its componentwise positive and negative parts and derive separate finite-time rates for the two components. The resulting certificates can assign a slower exponential envelope to the positive component than to the negative component. This rate separation provides indirect theoretical evidence for max-induced overestimation: positive errors can be amplified through the maximization step, whereas negative errors admit a sharper comparison with an optimal-policy system. The separation is a difference between upper bounds, so it need not hold for every realized Q-learning trajectory. Nevertheless, we construct examples in which the predicted asymmetry appears in the actual trajectory. The analysis gives deterministic and stochastic constant-step-size bounds and clarifies how overestimation enters the switching-system dynamics of Q-learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。