arXiv:2502.05706stat.MLcs.LG2025-02

首次给出非线性TD(0)在多项式混合数据下的有限样本收敛分析。

Convergence of TD(0) under Polynomial Mixing with Nonlinear Function Approximation

  • 提出新型离散耦合方法,突破几何遍历性限制。
  • 在β>1、η∈(1/2,1]条件下,收敛速度达t⁻ᵇᐟ²和t⁻ʸᵞ阶。
  • 适用于非平稳初始化,无需采样分块或自适应步长。

时序差分学习(TD(0))是强化学习的基础,但其在非独立同分布数据与非线性函数近似下的有限样本行为仍不清楚。本文首次提供在多项式混合马尔可夫数据下对原始TD(0)的高概率、有限样本分析,仅假设霍尔德连续性和有界广义梯度。对于混合指数β>1、霍尔德指数γ、步长衰减率η∈(1/2,1],我们证明:以高概率,经过t=𝒪(1/ε²)次迭代后,‖θₜ - θ*‖ ≤ C(β, γ, η) t⁻ᵇᐟ² + C'(γ, η) t⁻ʸᵞ。该界匹配已知的i.i.d.情形,并在初始非平稳时仍成立。核心在于一种新型离散时间耦合技术,绕过几何遍历性要求,首次实现非线性TD(0)在现实混合条件下的保证。

原文摘要 · Abstract (English)

Temporal Difference Learning (TD(0)) is fundamental in reinforcement learning, yet its finite-sample behavior under non-i.i.d. data and nonlinear approximation remains unknown. We provide the first high-probability, finite-sample analysis of vanilla TD(0) on polynomially mixing Markov data, assuming only Holder continuity and bounded generalized gradients. This breaks with previous work, which often requires subsampling, projections, or instance-dependent step-sizes. Concretely, for mixing exponent $β> 1$, Holder continuity exponent $γ$, and step-size decay rate $η\in (1/2, 1]$, we show that, with high probability, \[ \| θ_t - θ^* \| \leq C(β, γ, η)\, t^{-β/2} + C'(γ, η)\, t^{-ηγ} \] after $t = \mathcal{O}(1/\varepsilon^2)$ iterations. These bounds match the known i.i.d. rates and hold even when initialization is nonstationary. Central to our proof is a novel discrete-time coupling that bypasses geometric ergodicity, yielding the first such guarantee for nonlinear TD(0) under realistic mixing.

强化学习TD学习收敛分析非线性逼近

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。