arXiv:2608.11480eess.SYcs.LG2026-08

用前向轨迹引导采样,让神经网络更高效求解安全控制问题。

Forward Trajectory Steering for Hamilton-Jacobi Reachability Analysis

论文配图:Forward Trajectory Steering for Hamilton-Jacobi Reachability Analysis
图 1 · 摘自论文原文
  • 基于前向轨迹生成自适应采样点,提升训练效率。
  • 在多个基准测试中相对误差更低,性能优于主流方法。
  • 无需多阶段训练或额外监督,适合工程落地应用。

哈密顿-雅可比(HJ)可达性为动态系统的安全控制提供了严格的数学框架,但其在高维空间中求解哈密顿-雅可比-伊斯阿克斯变分不等式偏微分方程的计算复杂度限制了实际应用。物理信息神经网络(PINNs)作为传统网格求解器的替代方案受到关注,但其性能对采样点分布高度敏感。现有基于PINNs的HJ可达性求解器需复杂的训练流程和辅助监督以获得准确的安全值函数。本文提出STEER2REACH(S2R),一种仅需在标准PINNs训练基础上最小修改的求解器。S2R的核心创新在于:通过结合当前值函数诱导的最优控制与扰动信号,并注入随机探索噪声,构建轻量级、低开销的自适应采样分布,引导前向轨迹生成采样点。实验表明,尽管方法简单,S2R在多个可达性基准测试中达到竞争力甚至更优的安全性能,且相对L2误差显著降低,同时无需多阶段训练或模型预测控制(MPC)监督。

原文摘要 · Abstract (English)

Hamilton-Jacobi (HJ) reachability provides a mathematically rigorous framework for safe control of dynamical systems, but its practical application is bottlenecked by the computational complexity of solving Hamilton-Jacobi-Isaacs variational inequality PDEs in high dimensions. Physics-informed neural networks (PINNs) have recently emerged as a promising alternative to classical mesh-based solvers, yet their performance is highly sensitive to the choice of collocation sampling. In order to learn accurate safety value functions, existing PINNs-based HJ reachability solvers must rely on complex training pipelines and auxiliary supervision. In this work, we propose STEER2REACH (S2R), a PINNs-based HJ reachability solver that requires minimal modification on top of standard PINNs training. S2R's key contribution is a lightweight, low-overhead adaptive collocation sampling distribution constructed by steering forward trajectories using a combination of the optimal control and disturbance signals induced by the current value function, with injected stochastic exploration noise. We demonstrate that despite its simplicity, S2R achieves competitive--and in some cases improved--performance on safety metrics while reducing relative L2 error across a range of reachability benchmarks compared with SoTA MPC-guided HJ reachability solvers, all without requiring multi-stage training or MPC-based supervision.

安全控制神经网络可达性分析强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。