提出相对TD学习的稳定性分析,解决高折扣因子下的收敛难题。
Stability and Sensitivity Analysis of Relative Temporal-Difference Learning: Extended Version
- 基于线性函数逼近,分析相对TD学习的稳定条件
- 当基线为状态动作过程的经验分布时,算法对任意折扣因子均稳定
- 证明参数估计偏差和协方差在高折扣下仍保持有界,适合长期规划场景
相对时序差分(Relative TD)学习通过从时序差分更新中减去一个基准值,缓解了当折扣因子趋近于1时传统TD方法收敛缓慢的问题。尽管该思想已在表格设置中被研究,但在函数逼近下的稳定性仍不明确。本文针对线性函数逼近下的相对TD学习建立了稳定性条件,指出基线分布的选择起关键作用。特别地,当基线取为状态动作过程的经验分布时,算法对任意非负基线权重和任意折扣因子均保持稳定。此外,本文还进行了参数估计的敏感性分析,刻画了渐近偏差与协方差。结果表明,随着折扣因子趋近于1,渐近协方差和渐近偏差均保持一致有界。
原文摘要 · Abstract (English)
Relative temporal-difference (TD) learning was introduced to mitigate the slow convergence of TD methods when the discount factor approaches one by subtracting a baseline from the temporal-difference update. While this idea has been studied in the tabular setting, stability guarantees with function approximation remain poorly understood. This paper analyzes relative TD learning with linear function approximation. We establish stability conditions for the algorithm and show that the choice of baseline distribution plays a central role. In particular, when the baseline is chosen as the empirical distribution of the state-action process, the algorithm is stable for any non-negative baseline weight and any discount factor. We also provide a sensitivity analysis of the resulting parameter estimates, characterizing both asymptotic bias and covariance. The asymptotic covariance and asymptotic bias are shown to remain uniformly bounded as the discount factor approaches one.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。