arXiv:2608.10896stat.MLcs.LG2026-08

提出无需调参的常步长TD学习推断方法,实现高效精准估计。

Self-Normalized Inference for Constant-Stepsize Temporal-Difference Learning under Markovian Sampling

论文配图:Self-Normalized Inference for Constant-Stepsize Temporal-Difference Learning under Markovian Sampling
图 1 · 摘自论文原文
  • 基于布朗桥自归一化,构建无需估计协方差的置信区间。
  • 单次遍历即可完成推断,内存不随轨迹增长。
  • 适用于策略评估,尤其适合长轨迹场景下的稳定推断。

常步长时序差分(TD)学习在策略评估中具有吸引力,但单个马尔可夫轨迹的推断需处理序列相关性与步长依赖的稳态目标。针对固定步长线性TD,本文建立了函数中心极限定理,其协方差保留了随机TD矩阵带来的乘性成分及稳态迭代误差。进一步推导出由同一轨迹驱动的并行里查德森-罗姆伯格(RR)递推的联合函数极限。通过布朗桥自归一化,获得预设状态值对比的渐近枢轴置信区域,无需估计长期协方差或选择带宽/批长。该方法支持单遍历实现,内存不随轨迹长度增长。在固定步长下,推断中心为RR稳态目标。此外,研究了跨周期步长不变、周期间递减的时域索引设计,在显式依赖于RR的速率窗口下,残差RR目标偏移、乘性余项和初始化效应在根-n尺度上可忽略,从而实现对投影贝尔曼解的推断。在FrozenLake与Garnet上的实验验证了稳态目标覆盖性、RR目标校正效果及时域设计的有限样本表现。

原文摘要 · Abstract (English)

Constant-stepsize temporal-difference (TD) learning is attractive for policy evaluation, but inference from a single Markov trajectory must account for serial dependence and a stepsize-dependent stationary target. For fixed-stepsize linear TD, we establish a functional central limit theorem whose covariance retains the multiplicative component induced by the random TD matrix and the stationary iterate error. We then derive a joint functional limit for parallel Richardson--Romberg (RR) recursions driven by the same trajectory. A Brownian-bridge self-normalizer yields asymptotically pivotal confidence regions for prespecified state-value contrasts without estimating the long-run covariance or selecting a bandwidth or batch length. For such a contrast, the procedure admits a one-pass implementation whose memory does not grow with the trajectory length. At a fixed stepsize, the inferential center is the RR stationary target. We also study horizon-indexed designs in which the stepsize remains constant within each run and decreases across longer horizons. Under an explicit RR-dependent rate window, the residual RR target shift, multiplicative remainder, and initialization effect are negligible at the root-$n$ scale, yielding inference for the projected Bellman solution. Experiments on FrozenLake and Garnet illustrate stationary-target coverage, RR target correction, and the finite-sample behavior of the horizon-indexed design.

强化学习推断方法时序差分置信区间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。