arXiv:2502.11968cs.LGcs.AI2025-02

揭示基于贝尔曼方程的强化学习在高维空间中的理论瓶颈

Theoretical Barriers in Bellman-Based Reinforcement Learning

  • 通过构造反例,证明采样状态无法有效利用结构化信息
  • 算法会忽略关键问题特征,导致学习效率低下
  • 该缺陷也适用于事后经验回放的可达性学习方法

针对高维空间中的强化学习算法,通常仅在采样状态上强制满足贝尔曼方程,并依赖泛化能力将知识传播至整个状态空间。本文首次形式化并揭示了这一普遍做法的根本局限性。我们构造出具有简单结构的反例问题,表明此类方法无法有效利用问题内在结构。研究发现,这些算法可能完全忽略关键问题信息,从而造成严重学习效率损失。进一步地,我们还将这一负面结果扩展至文献中另一类方法:基于事后经验回放的状态-状态可达性学习。

原文摘要 · Abstract (English)

Reinforcement Learning algorithms designed for high-dimensional spaces often enforce the Bellman equation on a sampled subset of states, relying on generalization to propagate knowledge across the state space. In this paper, we identify and formalize a fundamental limitation of this common approach. Specifically, we construct counterexample problems with a simple structure that this approach fails to exploit. Our findings reveal that such algorithms can neglect critical information about the problems, leading to inefficiencies. Furthermore, we extend this negative result to another approach from the literature: Hindsight Experience Replay learning state-to-state reachability.

强化学习贝尔曼方程理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。