arXiv:2606.18183stat.MLcs.LG2026-06

用随机微分方程揭示线性TD学习的误差下限成因

A Diffusion Approximation for Temporal-Difference Learning with Linear Features under Markovian Noise

论文配图:A Diffusion Approximation for Temporal-Difference Learning with Linear Features under Markovian Noise
图 1 · 摘自论文原文
  • 构建马尔可夫噪声下的线性TD(0)随机微分方程模型
  • 发现误差下限由采样协方差与投影贝尔曼算子几何共同决定
  • 适合研究强化学习收敛性与稳定性的人群阅读

带线性函数近似的时序差分(TD)学习是策略评估的核心方法。其经典连续时间描述为常微分方程(ODE),仅捕捉渐近均值动态,忽略决定误差下限的随机波动。本文提出在马尔可夫噪声下对线性TD(0)的随机微分方程(SDE)近似。该模型区分了受投影贝尔曼算子控制的收缩动力学与马尔可夫采样的影响。结果表明,常步长下的误差下限源于马尔可夫长期协方差与投影贝尔曼算子收缩几何的相互作用。

原文摘要 · Abstract (English)

Temporal difference (TD) learning with linear function approximation is a core method for policy evaluation. Its classical continuous-time description is an ordinary differential equation (ODE), which captures the asymptotic mean dynamics but neglects stochastic fluctuations determining the error floor. We introduce a stochastic differential equation (SDE) approximation for linear TD(0) under Markovian noise. The resulting model distinguishes the contraction dynamics governed by the projected Bellman operator from the influence of Markovian sampling. As a consequence, the model explains the constant-stepsize error floor through the interaction between Markovian long-run covariance and the contraction geometry of the projected Bellman operator.

强化学习TD学习随机微分方程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。