arXiv:2606.15311eess.SYcs.RO2026-06

用前瞻安全集提升紧急避障的稳定性与成功率。

Hamilton-Jacobi Reachability-Based Safe Reinforcement Learning for Emergency Collision Avoidance

论文配图:Hamilton-Jacobi Reachability-Based Safe Reinforcement Learning for Emergency Collision Avoidance
图 1 · 摘自论文原文
  • 基于哈密顿-雅可比可达性构建未来时域安全约束
  • 实车与仿真测试中避障成功率更高、轨迹更平滑
  • 适合高动态极限驾驶场景下的安全强化学习

极端驾驶条件下紧急避障需同时考虑障碍物接近度与车辆动力学稳定性,现有方法多依赖瞬时或局部安全评估。本文提出一种由哈密顿-雅可比(HJ)可达性驱动的安全强化学习框架,通过融合几何碰撞裕度与底盘稳定性限制,构建统一的有符号安全函数,并经可达性分析扩展为有限时域运动安全集,表征未来状态演化中能否保持安全。为实现高效计算,该安全集基于离线极端驾驶数据近似,减轻网格化HJ求解器的负担。学习得到的安全集作为连续安全代价嵌入约束马尔可夫决策过程,采用PID-Lagrangian策略优化方案自适应调节拉格朗日乘子以强化安全约束。在低附着路面避障场景的仿真与实车实验表明,该方法相比基线方法具有更高的目标达成率、更平滑的避障动作,且维持更大的统一安全裕度。

原文摘要 · Abstract (English)

Emergency collision avoidance under extreme driving conditions demands safety-critical control that accounts for both obstacle proximity and vehicle dynamic stability over a future time horizon, yet existing methods often rely on instantaneous or local safety evaluations. This paper proposes a safe reinforcement learning framework guided by a Hamilton-Jacobi (HJ) reachability based motion safety set that provides forward-looking safety supervision for constrained policy optimization. Specifically, a unified signed safety function is formulated by combining geometric collision margins and chassis stability limits, and is then extended through reachability analysis into a finite-horizon motion safety set that characterizes whether safety can be maintained under future vehicle state evolution. To enable practical computation, the motion safety set is approximated from offline extreme driving data, mitigating the computational burden of grid-based HJ solvers. The learned motion safety set is then embedded as a continuous safety cost into a constrained Markov decision process, and a PID-Lagrangian policy optimization scheme is employed to adaptively regulate the Lagrange multiplier for safety constraint enforcement. Simulation and real-vehicle experiments on low-adhesion obstacle-avoidance scenarios demonstrate that the proposed method achieves higher goal-reaching rates, produces smoother avoidance maneuvers, and maintains larger unified safety margins than baseline methods.

安全强化学习避障控制可达性分析实时控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。