arXiv:2608.19643cs.LGstat.ML2026-08

纠正强化学习中折扣最小二乘估计的集中不等式错误,提出正确边界。

Time-Uniform Self-Normalized Concentration for Discounted Least Squares: Limits and Corrections

  • 发现原方法因混合分布不一致导致时间统一界限失效
  • 证明在大时间下边界至少为 $R\sqrt{\log(T/δ)}$ 的量级
  • 给出固定时间与无限时域的修正方法,适用于严谨分析

自归一化集中不等式是强化学习和多臂赌博机分析中的标准工具。一种广泛使用的加权扩展声称在非平稳问题中对折扣最小二乘估计具有类似的时间统一保证。然而,一个固定参数的标量高斯反例表明,所声称的有界半径会以概率1被超越。对于固定的折扣和正则化参数,我们进一步证明:当 $δ≤1/2$ 且 $T/δ$ 足够大时,任何在给定条件子高斯模型类上对所有时间统一有效的确定性即时边界,在时间 $T$ 前必须至少达到 $R\sqrt{\log(T/δ)}$ 的阶;对于非递减边界,该阶数在时间 $T$ 处必须成立。我们识别出证明中的错误:不同终止时间使用不同的高斯混合分布,因此固定时间的混合物不构成一个超级鞅,停止时间论证无法修复这一缺陷。最后,我们证明加权不等式在每个固定确定时间仍有效,并给出了有限和无限时域的修正方法,讨论了其对下游分析的影响。

原文摘要 · Abstract (English)

Self-normalized concentration inequalities are standard tools in bandit and reinforcement-learning analyses. A widely used weighted extension claims an analogous time-uniform guarantee for discounted least-squares estimators in non-stationary problems. A simple scalar Gaussian counterexample with a fixed parameter shows that the claimed bounded radius is crossed with probability one. For fixed discount and regularization parameters, we further show that, when $δ\leq1/2$ and $T/δ$ is sufficiently large, any deterministic anytime boundary valid uniformly over the stated conditionally sub-Gaussian model class must be at least of order $R\sqrt{\log(T/δ)}$ at some time by horizon $T$; for nondecreasing boundaries, this order is required at time $T$. We identify the proof error: different terminal times use different Gaussian mixing distributions, so the fixed-time mixtures do not form one supermartingale, and the stopping-time argument does not repair this failure. Finally, we show that the weighted inequality remains valid at each fixed deterministic time, give valid finite- and infinite-horizon corrections, and discuss consequences for downstream analyses.

强化学习集中不等式理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。