提出在线统计推断方法,高效实现分布强化学习中的分位数时序差分分析。
Online Inference for Quantile Temporal Difference Learning in Distributional Reinforcement Learning
- 基于随机缩放构造渐近枢轴统计量,利用完整迭代路径信息进行推断。
- 证明同步与异步分位数时序差分算法的均值迭代弱收敛于缩放布朗运动。
- 可在线计算,无需存储全部迭代轨迹,显著降低内存开销,适合实时场景。
本文研究分布强化学习中分位数时序差分学习(QTD)的统计推断问题。在生成模型假设下,我们首先建立了同步与异步QTD的泛函中心极限定理,表明QTD的平均迭代序列弱收敛于一个缩放的布朗运动。随后,我们提出了在线推断方法:基于随机缩放,利用整个QTD路径的信息构建渐近枢轴统计量,实现统计推断。该统计量可在线计算,无需存储完整的迭代轨迹,大幅降低内存需求,使分布强化学习中的高效统计推断成为可能。
原文摘要 · Abstract (English)
In this paper, we study how to perform statistical inference for quantile temporal difference learning (QTD) in distributional reinforcement learning. Assuming access to a generative model, we first establish functional central limit theorems for both synchronous and asynchronous QTD, which show that the averaged iterates of QTD converge weakly to a rescaled Brownian motion. We next provide online inference methods. Based on random scaling, the inference procedure constructs an asymptotically pivotal statistic for inference by using the information along the whole QTD path. Meanwhile, the proposed statistic can be computed online without storing the entire trajectory of QTD iterates. This substantially reduces the memory requirement and enables efficient statistical inference in distributional reinforcement learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。