arXiv:2604.13325cs.ROcs.SY2026-04

用最优控制理论指导采样,让安全过滤器更快更准地学会避险。

Boundary Sampling to Learn Predictive Safety Filters via Pontryagin's Maximum Principle

论文配图:Boundary Sampling to Learn Predictive Safety Filters via Pontryagin's Maximum Principle
图 1 · 摘自论文原文
  • 基于庞特里亚金极小值原理,识别临界避险轨迹以指导数据采集。
  • 在赛车共享控制实验中,学习效率提升,失败率降低,3毫秒内完成推理。
  • 适合需要实时安全保障的高维自主系统,如自动驾驶车辆。

安全滤波器为自主系统施加安全约束提供了一种实用方法。尽管基于学习的工具可扩展至高维系统,但其性能依赖于包含可能导致约束违反状态的有信息量数据,而在复杂高维系统中高效采样此类数据颇具挑战。本文利用庞特里亚金极小值原理刻画几乎避免安全违规的轨迹,这些边界轨迹用于引导学习哈密顿-雅可比可达性的数据收集,使学习集中于安全临界状态附近,从而提升效率。所学的控制屏障值函数直接用于安全滤波。仿真与共享控制汽车竞速应用的实验验证表明,基于PMP的采样显著提升学习效率,实现更快收敛、更低失败率及更优安全集重构,计算耗时约3毫秒。

原文摘要 · Abstract (English)

Safety filters provide a practical approach for enforcing safety constraints in autonomous systems. While learning-based tools scale to high-dimensional systems, their performance depends on informative data that includes states likely to lead to constraint violation, which can be difficult to efficiently sample in complex, high-dimensional systems. In this work, we characterize trajectories that barely avoid safety violations using the Pontryagin Maximum Principle. These boundary trajectories are used to guide data collection for learned Hamilton-Jacobi Reachability, concentrating learning efforts near safety-critical states to improve efficiency. The learned Control Barrier Value Function is then used directly for safety filtering. Simulations and experimental validation on a shared-control automotive racing application demonstrate PMP sampling improves learning efficiency, yielding faster convergence, reduced failure rates, and improved safe set reconstruction, with wall times around 3ms.

安全过滤最优控制自主系统强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。