提出连续时间策略评估的统计保证,揭示误差权衡新机制。
Statistical guarantees for continuous-time policy evaluation: blessing of ellipticity and new tradeoffs
- 基于LSTD方法,在Sobolev范数下实现1/√T收敛率。
- 轨迹长度只需近线性于混合时间和基函数数即可达到最优速率。
- 揭示近似误差与统计误差间的新型权衡,适合强化学习理论研究者。
我们研究了在单条离散观测遍历轨迹下,对连续时间马尔可夫扩散过程的价值函数进行估计的问题。本文为最小二乘时序差分(LSTD)方法提供了非渐近统计保证,性能以一阶Sobolev范数衡量。具体而言,当轨迹长度为$T$时,估计器达到$O(1 / \ oot{2}\ oot{T})$的收敛速率;值得注意的是,该速率只要求$T$与扩散过程的混合时间及所用基函数数量近乎线性增长即可实现。我们方法的关键洞察在于,扩散过程中固有的椭圆性确保了即使有效时域趋于无穷,性能依然稳健。此外,我们证明了统计误差中的马尔可夫分量可由近似误差控制,而鞅分量随基函数数量增长更缓慢。通过精细平衡这两类误差,我们的分析揭示了近似误差与统计误差之间的新型权衡。
原文摘要 · Abstract (English)
We study the estimation of the value function for continuous-time Markov diffusion processes using a single, discretely observed ergodic trajectory. Our work provides non-asymptotic statistical guarantees for the least-squares temporal-difference (LSTD) method, with performance measured in the first-order Sobolev norm. Specifically, the estimator attains an $O(1 / \sqrt{T})$ convergence rate when using a trajectory of length $T$; notably, this rate is achieved as long as $T$ scales nearly linearly with both the mixing time of the diffusion and the number of basis functions employed. A key insight of our approach is that the ellipticity inherent in the diffusion process ensures robust performance even as the effective horizon diverges to infinity. Moreover, we demonstrate that the Markovian component of the statistical error can be controlled by the approximation error, while the martingale component grows at a slower rate relative to the number of basis functions. By carefully balancing these two sources of error, our analysis reveals novel trade-offs between approximation and statistical errors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。