用隐式哈密顿量训练高维最优控制,突破传统方法限制。
End-to-End Training of High-Dimensional Optimal Control with Implicit Hamiltonians via Jacobian-Free Backpropagation
- 直接参数化价值函数,利用梯度与控制律的理论关联。
- 通过无雅可比反向传播实现高效训练,支持高维系统。
- 适用于航天器再入等含隐式哈密顿量的实际问题。
基于神经网络参数化价值函数的方法在哈密顿量具有显式表达时,已成功逼近高维最优反馈控制器。然而,许多实际问题(如航天器再入、自行车动力学)涉及无法显式表达的隐式哈密顿量,限制了现有方法的应用。本文提出一种端到端的隐式深度学习方法,直接参数化价值函数以学习最优控制律。该方法通过利用最优控制与价值函数梯度之间的基本关系,强制模型遵守物理规律,这是庞特里亚金最大值原理与动态规划之间联系的直接结果。采用无雅可比反向传播(JFB),克服轨迹优化中的时间耦合问题,实现高效训练。我们证明了JFB能生成最优控制目标的下降方向,并在多个包含隐式哈密顿量的场景中实验验证,该方法有效学习出高维反馈控制器,而现有方法无法处理此类问题。
原文摘要 · Abstract (English)
Neural network approaches that parameterize value functions have succeeded in approximating high-dimensional optimal feedback controllers when the Hamiltonian admits explicit formulas. However, many practical problems, such as the space shuttle reentry problem and bicycle dynamics, among others, may involve implicit Hamiltonians that do not admit explicit formulas, limiting the applicability of existing methods. Rather than directly parameterizing controls, which does not leverage the Hamiltonian's underlying structure, we propose an end-to-end implicit deep learning approach that directly parameterizes the value function to learn optimal control laws. Our method enforces physical principles by ensuring trained networks adhere to the control laws by exploiting the fundamental relationship between the optimal control and the value function's gradient; this is a direct consequence of the connection between Pontryagin's Maximum Principle and dynamic programming. Using Jacobian-Free Backpropagation (JFB), we achieve efficient training despite temporal coupling in trajectory optimization. We show that JFB produces descent directions for the optimal control objective and experimentally demonstrate that our approach effectively learns high-dimensional feedback controllers across multiple scenarios involving implicit Hamiltonians, which existing methods cannot address.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。