揭示TD学习方差机制,提出用控制变量降方差的新方法。
On the Variance of Temporal Difference Learning and its Reduction Using Control Variates

- 通过分段设置分析方差来源,发现其本质是多轨迹独立聚合。
- TD方差在渐近下低于蒙特卡洛估计,短时程更新更优。
- 提出改进的优势函数估计法,大样本下方差更小,适合强化学习优化。
我们分析了基于表格表示的分段设定下时序差分(TD)学习的方差,发现其降低方差的关键机制在于有效聚合更多独立轨迹。基于此,我们证明:(1) TD的方差在渐近意义下被蒙特卡洛(MC)估计器上界控制;(2) 在固定样本数下,较短时程的更新具有更低方差。此外,我们表明直接优势估计(DAE)可视为一种回归调整型控制变量,在大样本极限下比TD具有更紧的方差上界。最后,我们在精心设计的环境中数值验证了这些估计器的行为表现。
原文摘要 · Abstract (English)
We analyze the variance of temporal difference (TD) learning using the phased setting with tabular representation, and show that one of the mechanisms behind its ability to reduce variance is by effectively aggregating over a larger number of independent trajectories. Based on this insight, we demonstrate that (1) the variance of TD is asymptotically bounded from above by Monte Carlo (MC) estimators, and (2) shorter horizon updates incurs less variance for a fixed number of samples. Beyond TD, we show that Direct Advantage Estimation (DAE), a method for estimating the advantage function, can be seen as a type of regression-adjusted control variate, which achieves a tighter bound on the variance compared to TD in the large-sample limit. Finally, we numerically illustrate the behaviors of these estimators with carefully designed environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。