arXiv:2506.01052cs.LGmath.OC2025-06被引 3

无投影TD学习实现鲁棒收敛,速度达1/√T

A Robust $\widetilde{\mathcal{O}}(1/\sqrt{T})$ Rate for Unprojected TD Learning with Linear Function Approximation

  • 不加投影的TD(0)算法,仅需微调学习率
  • 期望误差收敛速度为‖θ*‖²₂/√T,含马尔可夫噪声
  • 首次在无正则条件下实现该速率,适合强化学习研究者

我们研究了带线性函数近似的时序差分(TD)学习的有限时间收敛性,这是强化学习的核心方法。关注的是“鲁棒”设定,即收敛性不依赖于潜在函数的最小曲率。尽管已有工作在此设定下建立了收敛保证,但通常依赖于人为假设:每次迭代都投影到有界集。这一条件的去除被Bhandari等(COLT'18)列为开放问题,并推测需要额外的“正则性条件”。本文证明,简单的无投影TD(0)在期望下仍以$ ilde{/mathcal{O}}ig( rac{ orm{ heta^*}^2_2}{ oot{T}}ig)$的速率收敛,即使存在马尔可夫噪声。我们无需额外正则条件,仅需对学习率进行微小的polylog修正。分析揭示了TD更新的新自约束性质,并利用其保证迭代值有界。

原文摘要 · Abstract (English)

We investigate the finite-time convergence properties of Temporal Difference (TD) learning with linear function approximation, a cornerstone of reinforcement learning. We are interested in the so-called ``robust'' setting, where the convergence guarantee does not depend on the potential function's minimal curvature. While prior work has established convergence guarantees in this setting, these results typically rely on the artificial assumption that each iterate is projected onto a bounded set. Removing such a condition was left as an open problem by Bhandari et al. (COLT'18), hypothesizing the need for additional ``regularity conditions''. In this paper, we show that the simple unprojected TD(0) converges with a rate of $\widetilde{\mathcal{O}}\left(\frac{\|θ^*\|^2_2}{\sqrt{T}}\right)$ in expectation, even in the presence of Markovian noise. We do not require an additional regularity condition, but only a minor polylog correction to the learning rate. Our analysis reveals a novel self-bounding property of the TD updates and exploits it to guarantee bounded iterates.

强化学习TD学习收敛性线性逼近

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。