arXiv:2606.05967stat.MLcs.LG2026-06中稿 · AISTATS 2026

提出TD(0)新收敛率,快且对病态不敏感,适合稳定强化学习训练

Fast and Robust Convergence Rate for TD(0) with Linear Function Approximation, Universal Learning Steps and I.I.D. Samples

论文配图:Fast and Robust Convergence Rate for TD(0) with Linear Function Approximation, Universal Learning Steps and I.I.D. Samples
图 1 · 摘自论文原文
  • 用Polyak-Juditsky平均法+固定学习率,实现1/k阶收敛
  • 均方误差下降速度达1/k,不受协方差最小特征值影响
  • 适用于样本独立同分布场景,对模型病态鲁棒

本文研究了在独立同分布样本下,采用线性函数逼近的TD(0)方法的有限时间行为。考虑常数学习率与Polyak-Juditsky平均法,建立了新的均方误差(MSE)收敛率:(i) 快速性——迭代次数k的依赖为最优的1/k阶;(ii) 鲁棒性——仅依赖初始误差和模型无关常数,不依赖线性参数化未中心协方差矩阵的最小特征值;(iii) 紧致性——乘性常数小于11。该结果优于现有所有TD(0)文献中的O(1/k)率。此外,引入了PCTD(0)变体,在马尔可夫链强混合假设下具有更优收敛性质。

原文摘要 · Abstract (English)

In this paper, we study the finite-time behavior of the TD(0) temporal-difference method with linear function approximation (LFA). We consider on-policy independent and identically distributed (i.i.d.) samples, a constant learning step, and the Polyak-Juditsky averaging method. We establish a new convergence rate, for the Mean-Square Error (MSE) on the approximated function, that is (i) fast in the sense that it admits an optimal dependency in the number of iterations k (i.e., of order 1/k), (ii) robust to ill-conditioning: it only depends on an initial error and modelindependent constants and (iii) sharp up to a multiplicative constant lower than 11. In particular, it does not depend on the smallest eigenvalue of the uncentered covariance matrix of the linear parametrization, unlike all pre-existing O(1/k) rates in the TD(0) literature. We also introduce PCTD(0), a variant of TD(0), which benefits from better convergence properties under an additional assumption of strong mixing on the Markov Chain.

强化学习TD(0)收敛率函数逼近

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。