让黑箱混合系统在不确定动态下仍能严格满足安全约束。
Learning Control Policies to Provably Satisfy Hard Affine Constraints for Black-Box Hybrid Dynamical Systems
- 用仿射策略在边界附近保持排斥性,确保轨迹不越界。
- 在状态跳变前设置第二重排斥区,防止跳变后违规。
- 比现有方法更安全,适合高风险控制场景。
黑箱混合动力系统的安全性保障极具挑战,因其存在瞬时状态跳变和未知的非线性动态。现有严格安全约束满足方法如控制屏障函数(CBFs)和可达性分析依赖系统动力学的显式知识,而安全强化学习通常依赖已知动力学或通过奖励设计间接抑制违规行为。本文旨在为具有仿射重置映射的黑箱混合动力系统,学习可证明在闭环中满足仿射状态约束的强化学习策略。核心思想是强制策略在未知非线性动态的约束边界附近呈仿射且排斥性,从而保证轨迹不会违反约束。进一步地,通过在重置前引入第二个排斥仿射区域,防止因碰撞或重置映射导致的状态跳变引发约束违反。我们推导了确保闭环安全的充分条件。在受约束摆和击球器等混合动力系统环境中,与最先进的奖励塑形及学习型CBF方法对比显示,本方法所学策略质量更高且始终满足安全约束。
原文摘要 · Abstract (English)
Ensuring safety for black-box hybrid dynamical systems presents significant challenges due to their instantaneous state jumps and unknown explicit nonlinear dynamics. Existing solutions for strict safety constraint satisfaction, like control barrier functions (CBFs) and reachability analysis, rely on direct knowledge of the dynamics. Similarly, safe reinforcement learning (RL) approaches often rely on known system dynamics or merely discourage safety violations through reward shaping. In this work, we want to learn RL policies which provably satisfy affine state constraints in closed loop for black-box hybrid dynamical systems with affine reset maps. Our key insight is forcing the RL policy to be affine and repulsive near the constraint boundaries for the unknown nonlinear dynamics of the system, providing guarantees that the trajectories will not violate the constraint. We further account for constraint violation due to instantaneous state jumps that occur due to impacts or reset maps in the hybrid system by introducing a second repulsive affine region before the reset that prevents post-reset states from violating the constraint. We derive sufficient conditions under which these policies satisfy safety constraints in closed loop. We also compare our approach with state-of-the-art reward shaping and learned-CBF methods on hybrid dynamical systems like the constrained pendulum and paddle juggler environments. In both scenarios, we show that our methodology learns higher quality policies while always satisfying the safety constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。