用新型屏障函数让飞行器和自动驾驶零违规地安全运行。
Barrier Function Overrides For Non-Convex Fixed Wing Flight Control and Self-Driving Cars
- 为非凸系统设计离散时间屏障函数,实现安全控制
- 近似解法性能接近基线强化学习,且无安全违规
- 适合高风险机器人控制场景,兼顾安全与效率
强化学习(RL)已显著提升机器人系统性能,但其随机探索机制在安全关键系统中带来挑战。屏障函数可通过近似保留RL控制输入的同时强制满足安全约束来解决该问题。然而,当系统动力学对控制输入非凸或时间离散时(如多数RL训练场景),该安全覆盖操作可能计算不可行。本文针对固定翼飞机和自适应巡航控制下的自动驾驶车辆变道任务,研究非凸、离散时间系统的新型屏障函数。尽管在线最优覆盖求解一般不可行,我们探索了近似解法,发现其性能可媲美基线强化学习方法,且实现零安全违规。尤其值得注意的是,即使不尝试求解最优覆盖,性能仍具备竞争力。文中讨论了近似方案在性能与计算可行性间的权衡。
原文摘要 · Abstract (English)
Reinforcement Learning (RL) has enabled vast performance improvements for robotics systems. To achieve these results though, the agent often must randomly explore the environment, which for safety critical systems presents a significant challenge. Barrier functions can solve this challenge by enabling an override that approximates the RL control input as closely as possible without violating a safety constraint. Unfortunately, this override can be computationally intractable in cases where the dynamics are not convex in the control input or when time is discrete, as is often the case when training RL systems. We therefore consider these cases, developing novel barrier functions for two non-convex systems (fixed wing aircraft and self-driving cars performing lane merging with adaptive cruise control) in discrete time. Although solving for an online and optimal override is in general intractable when the dynamics are nonconvex in the control input, we investigate approximate solutions, finding that these approximations enable performance commensurate with baseline RL methods with zero safety violations. In particular, even without attempting to solve for the optimal override at all, performance is still competitive with baseline RL performance. We discuss the tradeoffs of the approximate override solutions including performance and computational tractability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。