Bellman方程在连续状态空间中解不唯一,易导致控制失效。
Is Bellman Equation Enough for Learning Control?
- 发现线性系统下贝尔曼方程至少有组合数个解
- 仅一个解能同时保证最优策略与闭环稳定
- 提出正定神经架构确保收敛到稳定解
贝尔曼方程及其连续时间版本——哈密顿-雅可比-贝尔曼(HJB)方程,在强化学习与最优控制中是极值的必要条件。尽管在表格型设置中价值函数是贝尔曼方程的唯一解,我们证明在连续状态空间中该唯一性不成立。具体而言,对于线性动力系统,贝尔曼方程至少存在 $\binom{2n}{n}$ 个解,其中 $n$ 为状态维度。关键的是,仅有其中一个解能同时产生最优策略并维持闭环系统的稳定性。我们进一步揭示了基于价值的方法常见失败模式:由于可接受解与不可接受解之间呈指数级不平衡,模型可能收敛至不稳定解。最后,我们提出一种正定神经架构,通过构造保证收敛至稳定解。
原文摘要 · Abstract (English)
The Bellman equation and its continuous-time counterpart, the Hamilton-Jacobi-Bellman (HJB) equation, serve as necessary conditions for optimality in reinforcement learning and optimal control. While the value function is known to be the unique solution to the Bellman equation in tabular settings, we demonstrate that this uniqueness fails to hold in continuous state spaces. Specifically, for linear dynamical systems, we prove the Bellman equation admits at least $\binom{2n}{n}$ solutions, where $n$ is the state dimension. Crucially, only one of these solutions yields both an optimal policy and a stable closed-loop system. We then demonstrate a common failure mode in value-based methods: convergence to unstable solutions due to the exponential imbalance between admissible and inadmissible solutions. Finally, we introduce a positive-definite neural architecture that guarantees convergence to the stable solution by construction to address this issue.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。