揭示异步平均Q学习的渐近正态性,为算法稳定性提供理论支撑。
Central Limit Theorems for Asynchronous Averaged Q-Learning
- 基于Polyak-Ruppert平均,建立非渐近中心极限定理
- 收敛速率显式依赖迭代次数、状态动作空间大小等
- 适用于分析强化学习中异步更新的统计行为
本文建立了在异步更新下Polyak-Ruppert平均Q-learning的中心极限定理。我们证明了一个非渐近中心极限定理,其在Wasserstein距离下的收敛速率明确反映了迭代次数、状态-动作空间规模、折扣因子以及探索质量的影响。此外,我们推导出一个泛函中心极限定理,表明部分和过程弱收敛于布朗运动。
原文摘要 · Abstract (English)
This paper establishes central limit theorems for Polyak-Ruppert averaged Q-learning under asynchronous updates. We prove a non-asymptotic central limit theorem, where the convergence rate in Wasserstein distance explicitly reflects the dependence on the number of iterations, state-action space size, the discount factor, and the quality of exploration. In addition, we derive a functional central limit theorem, showing that the partial-sum process converges weakly to a Brownian motion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。