arXiv:2605.04100cs.AI2026-05

改进强化学习中不在线策略的稳定性与方差控制问题。

Regularized Centered Emphatic Temporal Difference Learning

  • 通过中心化修正提升算法投影几何特性。
  • 新方法在多个测试任务中避免了不稳定现象。
  • 适合研究稳定强化学习算法的学者参考。

不在线策略的时序差分(TD)学习在函数逼近下面临稳定性、投影几何和方差控制之间的结构性权衡。强调型TD(ETD)通过后续强调改善了不在线策略的投影几何,但后续迹可能具有高方差。本文通过贝尔曼误差中心化重新审视这一权衡。尽管中心化自然消除了TD误差中的常见漂移项,但直接的中心化强调扩展引入了辅助耦合,可能破坏ETD关键矩阵的正定性。为此,我们提出正则化强调型时序差分学习(RETD),保留后续迹并仅对辅助中心化递归进行正则化,相当于将耦合关键矩阵的右下块从1提升至1+c。我们推导出RETD的核心矩阵,在保守充分正则化条件下证明其收敛性,并在诊断性线性不在线策略预测任务上评估该方法。实验表明,RETD避免了朴素中心化强调学习的不稳定性,保持了有利的强调几何特性,并在正则化参数c上表现出稳健的中间区间。

原文摘要 · Abstract (English)

Off-policy temporal-difference (TD) learning with function approximation faces a structural tradeoff among stability, projection geometry, and variance control. Emphatic TD (ETD) improves the off-policy projection geometry through follow-on emphasis, but the follow-on trace can have high variance. We revisit this tradeoff through Bellman-error centering. Although centering naturally removes a common drift term from TD errors, we show that a naive centered emphatic extension introduces an auxiliary coupling that can destroy the positive-definiteness of the ETD key matrix. We propose \emph{Regularized Emphatic Temporal-Difference Learning} (RETD), which preserves the follow-on trace and regularizes only the auxiliary centering recursion, corresponding to lifting the lower-right block of the coupled key matrix from \(1\) to \(1+c\). We derive the RETD core matrix, prove convergence under a conservative sufficient regularization condition, and evaluate the method on diagnostic linear off-policy prediction tasks. The experiments show that RETD avoids the instability of naive centered emphatic learning, preserves favorable emphatic geometry, and exhibits a robust intermediate regime for the regularization parameter \(c\) across the diagnostics.

强化学习时序差分稳定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。