用在线探索提升安全控制,无需备用控制器。
Learning Safe Control via On-the-Fly Bandit Exploration
- 结合控制屏障函数与高斯过程,动态引导数据采集。
- 在模型不确定性高时仍能保证闭环系统安全。
- 适合高不确定性环境下的安全强化学习应用。
在高度模型不确定性的安全控制任务中,机器学习方法常依赖模型误差界来设计鲁棒的安全过滤器。然而,当学习到的模型不确定性过高时,安全过滤器可能失效,导致无可行控制输入。现有方法通常假设存在安全备份控制器,而本文通过基于高斯过程的贝叶斯多臂老虎机算法,在线收集额外数据。将控制屏障函数与学习模型结合,构造可验证的安全证书;当出现不可行时,利用屏障函数指导探索,确保所采集数据有助于闭环系统安全。该方法在零均值先验动力学模型下仍可保证安全性,且无需备份控制器,据我们所知是首个实现此目标的学习型安全控制方法。
原文摘要 · Abstract (English)
Control tasks with safety requirements under high levels of model uncertainty are increasingly common. Machine learning techniques are frequently used to address such tasks, typically by leveraging model error bounds to specify robust constraint-based safety filters. However, if the learned model uncertainty is very high, the corresponding filters are potentially invalid, meaning no control input satisfies the constraints imposed by the safety filter. While most works address this issue by assuming some form of safe backup controller, ours tackles it by collecting additional data on the fly using a Gaussian process bandit-type algorithm. We combine a control barrier function with a learned model to specify a robust certificate that ensures safety if feasible. Whenever infeasibility occurs, we leverage the control barrier function to guide exploration, ensuring the collected data contributes toward the closed-loop system safety. By combining a safety filter with exploration in this manner, our method provably achieves safety in a setting that allows for a zero-mean prior dynamics model, without requiring a backup controller. To the best of our knowledge, it is the first safe learning-based control method that achieves this.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。