突破性证明线性TD学习在任意特征下几乎必然收敛。
Almost Sure Convergence of Linear Temporal Difference Learning with Arbitrary Features
- 无需特征线性无关假设,直接分析权重迭代的不变集
- 证明权重收敛到有界集合,估值结果几乎处处相同
- 适用于特征相关性强的实际场景,如高维离散表示
线性时间差分(TD)学习是强化学习中经典的预测算法。传统理论认为其几乎必然收敛至唯一解,但依赖特征线性独立的前提,这在实际中常不成立。本文首次在无需特征线性独立假设下证明线性TD的几乎必然收敛性。我们证明权重迭代收敛至有界集合,且该集合内权重导出的值估计几乎处处相同,并建立了权重迭代的局部稳定性概念。分析核心在于对线性TD均值微分方程的有界不变集提出新刻画。研究未引入额外假设或修改算法,具有强普适性。
原文摘要 · Abstract (English)
Temporal difference (TD) learning with linear function approximation (linear TD) is a classic and powerful prediction algorithm in reinforcement learning. While it is well-understood that linear TD converges almost surely to a unique point, this convergence traditionally requires the assumption that the features used by the approximator are linearly independent. However, this linear independence assumption does not hold in many practical scenarios. This work is the first to establish the almost sure convergence of linear TD without requiring linearly independent features. We prove that the weight iterates of linear TD converge to a bounded set, and that the value estimates derived from the weights in that set are the same almost everywhere. We also establish a notion of local stability of the weight iterates. Importantly, we do not impose assumptions tailored to feature dependence and do not modify the linear TD algorithm. Key to our analysis is a novel characterization of bounded invariant sets of the mean ODE of linear TD.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。