用行为感知矩阵替代传统修正项,提升离策略时序差分学习稳定性。
Behavior-Aware Auxiliary Corrections for Off-Policy Temporal-Difference Prediction

- 以行为贝尔曼矩阵替代原辅助协方差矩阵,实现更合理的几何修正。
- 在两状态反例、Baird反例等任务上验证了行为感知修正的有效性。
- 适合研究值函数近似中离策略学习稳定性的研究人员参考。
使用函数逼近的时序差分学习在离策略采样下可能不稳定。TDC通过辅助协方差修正稳定离策略TD,TDRC进一步在单尺度递归中正则化该修正。本文在线性预测设置下研究行为感知的辅助协方差几何替换,这是理解值函数近似特征空间动态的标准局部模型。首先用行为贝尔曼矩阵(A_μ)替代TDC的辅助矩阵(C),得到BA-TDC;再对同一行为感知方程进行正则化,获得BA-TDRC。该两步构造分离了行为感知几何与正则化的贡献。线性分析还为神经网络值函数近似中的辅助几何设计问题提供了可处理的模型,其中特征协方差与时间转移矩阵共同决定最后一层修正动态。我们给出了有限状态均值系统形式,证明在均值系统满足赫尔维茨稳定性条件下具有固定点保持和几乎必然收敛性,并通过精确线性误差递归的谱半径比较确定性均值速率。在两状态反例、Baird反例、Random Walk和Boyan Chain上的实验表明,行为感知替换本身在某些任务上已显著有益,但在更复杂设置中正则化是实现鲁棒性能的必要条件。
原文摘要 · Abstract (English)
Temporal-difference learning with function approximation can be unstable under off-policy sampling. TDC stabilizes off-policy TD through an auxiliary covariance correction, and TDRC further regularizes this correction in a single-timescale recursion. This paper studies a behavior-aware replacement of the auxiliary covariance geometry in the linear prediction setting, which is the standard local model for understanding the feature-space dynamics of value-function approximation. We first replace the TDC auxiliary matrix (C) by the behavior Bellman matrix (A_μ), yielding BA-TDC, and then regularize the same behavior-aware equation to obtain BA-TDRC. This two-step construction separates the contribution of behavior-aware geometry from the contribution of regularization. The linear analysis also provides a tractable model for an auxiliary-geometry design question that arises in neural-network value approximation, where feature covariances and temporal transition matrices jointly shape the last-layer correction dynamics. We give a finite-state mean-system formulation, prove fixed-point preservation and almost-sure convergence under a Hurwitz stability condition on the instantiated mean system, and compare deterministic mean rates through the spectral radius of the exact linear error recursion. Experiments on the two-state counterexample, Baird's counterexample, Random Walk, and Boyan Chain show that the behavior-aware replacement can be highly beneficial by itself on some tasks, but that regularization is necessary for robust performance across harder settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。