arXiv:2601.18840cs.LGcs.SY2026-01被引 1

提出控制任务的贝尔曼残差最小化新方法,解决收敛难题。

Bellman Residual Minimization for Control: Geometry, Stationarity, and Convergence

  • 直接最小化贝尔曼残差,避免动态规划迭代
  • 证明了该方法在函数逼近下的稳定收敛性
  • 为强化学习中的策略优化提供新思路

马尔可夫决策问题通常通过动态规划求解。另一种方法是贝尔曼残差最小化,即直接最小化平方贝尔曼残差目标函数。然而,与动态规划相比,该方法关注较少,主要因其在实践中效率较低,且难以推广到无模型设置(如强化学习)。尽管如此,贝尔曼残差最小化具有若干优势,例如在值函数函数逼近下收敛更稳定。虽然贝尔曼残差方法在策略评估中已广泛研究,但在策略优化(控制任务)方面的研究仍很稀少。本文首次为控制任务的贝尔曼残差最小化建立基础理论结果,推动其在策略优化中的应用。

原文摘要 · Abstract (English)

Markov decision problems are most commonly solved via dynamic programming. Another approach is Bellman residual minimization, which directly minimizes the squared Bellman residual objective function. However, compared to dynamic programming, this approach has received relatively less attention, mainly because it is often less efficient in practice and can be more difficult to extend to model-free settings such as reinforcement learning. Nonetheless, Bellman residual minimization has several advantages that make it worth investigating, such as more stable convergence with function approximation for value functions. While Bellman residual methods for policy evaluation have been widely studied, methods for policy optimization (control tasks) have been scarcely explored. In this paper, we establish foundational results for the control Bellman residual minimization for policy optimization.

强化学习贝尔曼残差策略优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。