arXiv:2604.04218stat.MLcs.LG2026-04

提出新型学习率,让Q-learning兼具快速收敛与稳定精度。

Sharp asymptotic theory for Q-learning with LDTZ learning rate and its generalization

  • 用幂律衰减学习率(PD2Z-ν)替代传统固定或多项式衰减。
  • 证明该方法可实现快速收敛且渐近无偏,理论误差更小。
  • 提供统计推断工具,适合需可信结果的强化学习研究者。

尽管Q-learning在策略确定中广受欢迎,但现有理论多聚焦于固定(η_t≡η)或多项式衰减(η_t = ηt^{-α})的学习率,前者存在持续偏差,后者收敛过慢。近期提出的线性衰减至零(LD2Z: η_{t,n}=η(1-t/n))在实践中表现优异,但其理论性质尚未充分研究。本文首先考虑一类广义的幂律衰减至零(PD2Z-ν: η_{t,n}=η(1-t/n)^ν),逐步推导出带PD2Z-ν的学习率下Q-learning的精确非渐近误差界,并由此构建新的尾部Polyak-Ruppert平均估计器的中心极限定理。此外,还首次给出Q-learning迭代部分和过程的时间统一高斯逼近(即强不变原理),支持基于自助法的统计推断。所有理论均通过大量数值实验验证。结果表明,LD2Z及一般PD2Z-ν能兼顾初始化时的快速下降与渐近收敛保证,解释了其经验成功,并为实际推断提供指导。

原文摘要 · Abstract (English)

Despite the sustained popularity of Q-learning as a practical tool for policy determination, a majority of relevant theoretical literature deals with either constant ($η_{t}\equiv η$) or polynomially decaying ($η_{t} = ηt^{-α}$) learning schedules. However, it is well known that these choices suffer from either persistent bias or prohibitively slow convergence. In contrast, the recently proposed linear decay to zero (\texttt{LD2Z}: $η_{t,n}=η(1-t/n)$) schedule has shown appreciable empirical performance, but its theoretical and statistical properties remain largely unexplored, especially in the Q-learning setting. We address this gap in the literature by first considering a general class of power-law decay to zero (\texttt{PD2Z}-$ν$: $η_{t,n}=η(1-t/n)^ν$). Proceeding step-by-step, we present a sharp non-asymptotic error bound for Q-learning with \texttt{PD2Z}-$ν$ schedule, which then is used to derive a central limit theory for a new \textit{tail} Polyak-Ruppert averaging estimator. Finally, we also provide a novel time-uniform Gaussian approximation (also known as \textit{strong invariance principle}) for the partial sum process of Q-learning iterates, which facilitates bootstrap-based inference. All our theoretical results are complemented by extensive numerical experiments. Beyond being new theoretical and statistical contributions to the Q-learning literature, our results definitively establish that \texttt{LD2Z} and in general \texttt{PD2Z}-$ν$ achieve a best-of-both-worlds property: they inherit the rapid decay from initialization (characteristic of constant step-sizes) while retaining the asymptotic convergence guarantees (characteristic of polynomially decaying schedules). This dual advantage explains the empirical success of \texttt{LD2Z} while providing practical guidelines for inference through our results.

强化学习Q-learning学习率统计推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。