arXiv:2602.00921math.OCcs.LG2026-02被引 4

提出无需雅可比矩阵的反向传播方法,解决隐式哈密顿系统下的高维最优控制问题。

On the Convergence of Jacobian-Free Backpropagation for Optimal Control Problems with Implicit Hamiltonians

  • 使用无雅可比反向传播(JFB)求解隐式哈密顿系统的最优控制。
  • 在随机小批量设置下证明了算法收敛到期望目标的驻点。
  • 验证了在多智能体消费、无人机群和自行车控制等高维场景中的可扩展性。

基于学习的价值函数方法在处理具有隐式哈密顿量的最优反馈控制时面临根本挑战,因为缺乏闭式最优控制律。近期工作引入了基于无雅可比反向传播(JFB)的隐式深度学习方法以应对该问题,但仅建立了样本级下降的保证。本文在随机小批量设定下建立了JFB的收敛性保证,表明其更新方向收敛至期望最优控制目标的驻点。进一步在更高维度的问题上展示了可扩展性,包括多智能体最优消费及基于群体的四旋翼与自行车控制。结果共同为在高维隐式哈密顿系统中使用JFB提供了理论依据与实证支持。

原文摘要 · Abstract (English)

Optimal feedback control with implicit Hamiltonians poses a fundamental challenge for learning-based value function methods due to the absence of closed-form optimal control laws. Recent work~\cite{gelphman2025end} introduced an implicit deep learning approach using Jacobian-Free Backpropagation (JFB) to address this setting, but only established sample-wise descent guarantees. In this paper, we establish convergence guarantees for JFB in the stochastic minibatch setting, showing that the resulting updates converge to stationary points of the expected optimal control objective. We further demonstrate scalability on substantially higher-dimensional problems, including multi-agent optimal consumption and swarm-based quadrotor and bicycle control. Together, our results provide both theoretical justification and empirical evidence for using JFB in high-dimensional optimal control with implicit Hamiltonians.

最优控制深度学习无雅可比高维系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。