arXiv:2601.20599cs.LGcs.AI2026-01

解决强化学习中梯度时序差分算法在奇异情况下的收敛难题

R-GTD: A Geometric Analysis of Gradient Temporal-Difference Learning in Singular Regimes

  • 通过重构目标函数引入正则化,使算法在特征交互矩阵奇异时仍能收敛
  • 理论证明算法在奇异情形下收敛到唯一解,且给出明确误差上界
  • 适用于需要稳定评估的离线策略强化学习场景

梯度时序差分(GTD)学习算法广泛用于带函数逼近的离线策略评估。然而,现有收敛性分析依赖于特征交互矩阵(FIM)非奇异这一严格假设。实际中FIM可能奇异,导致算法不稳定或性能下降。尽管已有工作通过正则化放松该假设,但其理论保证仍需其他限制条件。本文通过重构最小化均方投影贝尔曼误差的目标函数,提出正则化优化目标,自然导出一种称为R-GTD的新算法。该算法在FIM奇异时仍能保证收敛至唯一解。我们进行了几何分析,建立了理论收敛性保证和显式误差上界,并通过实验验证了方法的有效性。

原文摘要 · Abstract (English)

Gradient temporal-difference (GTD) learning algorithms are widely used for off-policy policy evaluation with function approximation. However, existing convergence analyses rely on the restrictive assumption that the so-called feature interaction matrix (FIM) is nonsingular. In practice, the FIM can become singular and leads to instability or degraded performance. While some prior works have applied regularization to relax the nonsingularity assumption, their theoretical guarantees inevitably rely on other restrictive conditions. In this paper, we propose a regularized optimization objective by reformulating the mean-square projected Bellman error minimization. This formulation naturally yields a regularized GTD algorithms, referred to as R-GTD, which guarantees convergence to a unique solution even when the FIM is singular. We conduct a geometric analysis to establish theoretical convergence guarantees and explicit error bounds for the proposed method, and validate its effectiveness through empirical experiments.

强化学习收敛性分析正则化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。