arXiv:2502.13822stat.MLcs.LG2025-02被引 8

为马尔可夫链诱导的鞅建立不确定性量化方法,用于强化学习中的时序差分算法。

Uncertainty quantification for Markov chain induced martingales with application to temporal difference learning

  • 基于马尔可夫链构建向量鞅的高维浓度不等式与Berry-Esseen界。
  • 给出TD学习的高概率一致性保证,逼近渐近方差仅差对数因子。
  • 首次获得TD估计器高斯逼近的$O(T^{-1/4}"log T)$收敛速率,适合关注统计推断的RL研究者。

我们建立了针对由马尔可夫链诱导的向量值鞅的新型通用高维浓度不等式和Berry-Esseen界。将这些结果应用于具有线性函数近似的时序差分(Temporal Difference, TD)学习算法,该方法在强化学习中广泛用于策略评估,获得了与渐近方差匹配至对数因子的尖锐高概率一致性保证。此外,我们建立了TD估计器高斯近似在凸距离下的$O(T^{-1/4}\ \log T)$分布收敛速率。所提出的鞅界具有广泛的适用性,对TD学习的分析为强化学习算法的统计推断提供了新见解,弥合了经典随机逼近理论与现代强化学习应用之间的差距。

原文摘要 · Abstract (English)

We establish novel and general high-dimensional concentration inequalities and Berry-Esseen bounds for vector-valued martingales induced by Markov chains. We apply these results to analyze the performance of the Temporal Difference (TD) learning algorithm with linear function approximations, a widely used method for policy evaluation in Reinforcement Learning (RL), obtaining a sharp high-probability consistency guarantee that matches the asymptotic variance up to logarithmic factors. Furthermore, we establish an $O(T^{-\frac{1}{4}}\log T)$ distributional convergence rate for the Gaussian approximation of the TD estimator, measured in convex distance. Our martingale bounds are of broad applicability, and our analysis of TD learning provides new insights into statistical inference for RL algorithms, bridging gaps between classical stochastic approximation theory and modern RL applications.

强化学习时序差分不确定性量化鞅分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。