arXiv:2410.16106stat.MLcs.LG2024-10被引 14

为强化学习中的价值函数估计提供可信赖的统计推断方法

Statistical Inference for Policy Evaluation with Temporal Difference Learning

  • 用Polyak-Ruppert平均改进TD学习的统计性质
  • 首次实现有限样本下置信区间的精确覆盖
  • 适合需要可靠推断的强化学习研究者

我们研究了在独立样本假设下,使用Polyak-Ruppert平均的时序差分(TD)学习在估计最优线性价值函数近似参数时的统计性质。提出了三项理论贡献:(i) 在凸集类上建立了更精细的高维Berry-Esseen界,收敛速度优于现有最佳结果;(ii) 提出并分析了一种新型、计算高效的在线插值协方差矩阵估计器;(iii) 推导出依赖于渐近方差的更紧致的高概率收敛保证,且适用条件弱于文献中常见设定。这些结果使构建具有保证的有限样本覆盖率的价值函数线性参数置信区域和联合置信区间成为可能。通过数值实验验证了理论结果的实用性。

原文摘要 · Abstract (English)

We investigate the statistical properties of Temporal Difference (TD) learning with Polyak-Ruppert averaging, arguably one of the most widely used algorithms in reinforcement learning, for the task of estimating the parameters of the optimal linear approximation to the value function. Assuming independent samples, we make three theoretical contributions that improve upon the current state-of-the-art results: (i) we establish refined high-dimensional Berry-Esseen bounds over the class of convex sets, achieving faster rates than the best known results, and (ii) we propose and analyze a novel, computationally efficient online plug-in estimator of the asymptotic covariance matrix; (iii) we derive sharper high probability convergence guarantees that depend explicitly on the asymptotic variance and hold under weaker conditions than those adopted in the literature. These results enable the construction of confidence regions and simultaneous confidence intervals for the linear parameters of the value function approximation, with guaranteed finite-sample coverage. We demonstrate the applicability of our theoretical findings through numerical experiments.

强化学习统计推断价值函数

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。