arXiv:2507.15455math.NAcs.AI2025-07被引 2

用神经网络迭代求解高维非凸博弈方程,精度稳定且无需凸性假设。

Solving nonconvex Hamilton--Jacobi--Isaacs equations with PINN-based policy iteration

  • 结合动态规划与物理信息神经网络,交替优化策略和控制
  • 二维路径规划误差低于0.01%,五维十维任务优于直接PINN方法
  • 适合机器人、金融、多智能体强化学习中的高维非凸控制问题

我们提出一种无网格的策略迭代框架,将经典动态规划与物理信息神经网络(PINNs)结合,用于求解高维非凸哈密顿-雅可比-伊斯阿克斯(HJI)方程,该方程源于随机微分博弈与鲁棒控制。方法在固定反馈策略下交替求解线性二阶偏微分方程,并通过自动微分进行逐点极小极大优化更新控制。在标准Lipschitz和一致椭圆性假设下,证明了值函数迭代序列局部一致收敛至HJI方程的唯一黏性解。分析建立了迭代序列的一致Lipschitz正则性,从而在无需哈密顿量凸性条件下实现可证明的稳定性和收敛性。数值实验表明该方法具有高精度与可扩展性:在含移动障碍物的二维随机路径规划博弈中,相对$L^2$误差低于10^{-2}%;在具有各向异性噪声的五维与十维出版商-订阅者微分博弈中,所提方法持续优于直接PINN求解器,得到更平滑的值函数和更低残差。结果表明,将PINNs与策略迭代结合是一种实用且理论可靠的高维非凸HJI方程求解方法,潜在应用于机器人、金融与多智能体强化学习。

原文摘要 · Abstract (English)

We propose a mesh-free policy iteration framework that combines classical dynamic programming with physics-informed neural networks (PINNs) to solve high-dimensional, nonconvex Hamilton--Jacobi--Isaacs (HJI) equations arising in stochastic differential games and robust control. The method alternates between solving linear second-order PDEs under fixed feedback policies and updating the controls via pointwise minimax optimization using automatic differentiation. Under standard Lipschitz and uniform ellipticity assumptions, we prove that the value function iterates converge locally uniformly to the unique viscosity solution of the HJI equation. The analysis establishes equi-Lipschitz regularity of the iterates, enabling provable stability and convergence without requiring convexity of the Hamiltonian. Numerical experiments demonstrate the accuracy and scalability of the method. In a two-dimensional stochastic path-planning game with a moving obstacle, our method matches finite-difference benchmarks with relative $L^2$-errors below %10^{-2}%. In five- and ten-dimensional publisher-subscriber differential games with anisotropic noise, the proposed approach consistently outperforms direct PINN solvers, yielding smoother value functions and lower residuals. Our results suggest that integrating PINNs with policy iteration is a practical and theoretically grounded method for solving high-dimensional, nonconvex HJI equations, with potential applications in robotics, finance, and multi-agent reinforcement learning.

非凸控制PINN博弈论高维求解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。