arXiv:2608.14408stat.MLcs.LG2026-08

提出在线分布时序差分学习的统计推断方法,支持对回报分布的精准分析。

Online Inference in Distributional Temporal-Difference Learning

  • 使用平均化估计器结合非参数分布TD学习
  • 证明根T误差收敛到高斯过程,支持自助法推断
  • 适用于方差、分位数等平滑与非平滑统计量

我们研究固定策略下回报分布函数的在线统计推断。通过单个马尔可夫轨迹的非参数分布时序差分学习估计回报分布。对于Polyak-Ruppert平均估计器,我们证明其根-T误差在Cramér空间中弱收敛于一个中心高斯随机元。同时证明,在给定观测轨迹条件下,自助法与原始平均之间的根-T差值也弱收敛到同一高斯极限。该结果为光滑统计函数(如方差、条件风险价值CVaR、预期短路、期望分位数)的推断提供了理论依据。对于非光滑统计函数,我们建立了在有限多个阈值的T^{-1/2}邻域内回报累积分布函数的局部渐近理论及其自助法版本。该理论使基于累积分布函数方程的非光滑统计量(如回报分位数)的推断成为可能。

原文摘要 · Abstract (English)

We study online statistical inference for functionals of the return distribution under a fixed policy. The return distribution is estimated by nonparametric distributional temporal-difference learning from a single Markov trajectory. For the Polyak--Ruppert averaged estimator, we prove that its root-$T$ error converges weakly to a centered Gaussian random element in Cramér space. We also prove that, conditionally on the observed trajectory, the root-$T$ difference between the bootstrap and original averages converges weakly to the same Gaussian limit. These results justify bootstrap inference for smooth statistical functionals, including variance, CVaR, expected shortfall, and expectiles. For nonsmooth statistical functionals, we develop a local asymptotic theory for the estimated return CDF over $T^{-1/2}$-neighborhoods of finitely many thresholds, together with its bootstrap analogue. This theory allows us to conduct inference for nonsmooth statistical functionals characterized by CDF equations, including return quantiles.

强化学习分布学习统计推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。