arXiv:2508.01718cs.LGcs.CE2025-08被引 3

无需边界数据,用神经网络求解高维哈密顿-雅可比-贝尔曼方程。

Physics-Informed Policy Iteration for High-Dimensional Hamilton--Jacobi--Bellman Equations: Interior Error Bounds without Boundary Data

  • 用无网格神经残差法求解策略评估的椭圆型偏微分方程。
  • 在无边界数据条件下实现指数衰减的内部误差界,误差由 $L^p$ 残差决定。
  • 适合高维随机控制问题,对复杂系统如倒立摆、四旋翼有效。

我们提出一种物理信息驱动的策略迭代方法,用于求解连续时间随机控制中的二阶哈密顿-雅可比-贝尔曼(HJB)方程。每个策略评估步骤对应一个线性椭圆型偏微分方程,采用无网格神经残差求解器近似;策略改进则基于代理梯度逐点执行。分析聚焦于有界域训练但无预设边界数据的情形。证明了对任意博雷尔马尔可夫策略,由PDE定义的评估问题是适定的;建立了贪婪映射在有界梯度范围内的利普希茨稳定性,并导出未解析边界信息的指数衰减估计。这些结果共同给出一个闭合的有限步内部误差界,其下界由连续 $L^p$ 残差($p>d$)和衰减振幅项决定。实验测量了估计中的各项量。在线性-二次测试平台上,梯度误差下界与新样本 $L^p$ 残差估计近乎线性相关;有限差分探测揭示训练缓冲区的益处。在相同架构与预算下,随着问题难度增加,固定策略训练明显比直接最小化非线性HJB残差更可靠。该方法在倒立摆、平面四旋翼及100维后验搜索问题上均取得有效反馈。非线性测试暴露两个仅凭采样残差无法察觉的局限:配点分布可能遗漏学习闭环的访问区域,且无边界数据评估可能导致贪婪更新依赖的梯度分量不确定。采用在线策略配点、滚动回放锚定评估和基于滚动回放的停止机制可提供有效的模型内防护。

原文摘要 · Abstract (English)

We develop a physics-informed policy-iteration method for stationary second-order Hamilton--Jacobi--Bellman equations arising in continuous-time stochastic control. Each policy-evaluation step is a linear elliptic PDE and is approximated by a mesh-free neural residual solver; policy improvement is then performed pointwise from the surrogate gradient. The analysis addresses bounded-domain training without prescribed boundary data. We prove well-posedness of the PDE-defined evaluation for every Borel Markov policy, establish Lipschitz stability of the greedy map on bounded gradient ranges, and derive an exponential attenuation estimate for unresolved boundary information. These ingredients yield a closed finite-step interior error bound whose floors are determined by a continuous $L^p$ residual, with a finite exponent $p>d$, and an attenuated amplitude term. The experiments measure the quantities in the estimate. On a linear--quadratic testbed with an exact reference, the gradient-error floor scales nearly linearly with a fresh-sample $L^p$ residual estimate, and finite-difference probes identify when a training buffer is beneficial. At matched architecture and budget, linear fixed-policy training becomes markedly more reliable than direct minimization of the nonlinear HJB residual as the tested problems become more difficult. The method also produces effective feedback on an inverted pendulum, a planar quadrotor, and a 100-dimensional posterior-seeking problem. These nonlinear tests expose two limitations not visible from sampled residuals alone: the collocation distribution may miss the region visited by the learned closed loop, and evaluation without boundary data may leave undetermined the gradient component on which the greedy update depends. On-policy collocation, rollout-anchored evaluation, and rollout-based stopping provide effective model-only safeguards.

控制神经网络偏微分方程强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。