在未知环境下,用概率方法保证控制安全并提升规划时长。
Conformal Reachability for Safe Control in Unknown Environments
- 结合置信预测与可达性分析,动态生成系统不确定性区间。
- 所学策略在7个场景中既保持高奖励又实现强可证明安全。
- 适合需要可靠安全性的自主系统,如无人机与自动驾驶。
可信自主系统的安全控制设计是核心挑战。然而,以往工作大多假设系统动力学已知或为确定性模型,或状态与动作空间有限,严重限制了应用范围。本文提出一种针对未知动力系统的概率验证框架,融合置信预测与可达性分析:在每一步利用置信预测获取未知动力学的合理不确定性区间,并通过可达性分析验证在此区间内安全性是否维持。进一步,我们开发了一种算法,用于训练控制策略,在优化预期奖励的同时最大化具有严格概率安全保障的规划时长。我们在七个涵盖四类任务(倒立摆、车道保持、无人机控制、安全导航)的安全控制设置中评估该方法,覆盖仿射与非线性安全约束。实验表明,所学策略在保持高平均奖励的同时,实现了最强的可证明安全性保障。
原文摘要 · Abstract (English)
Designing provably safe control is a core problem in trustworthy autonomy. However, most prior work in this regard assumes either that the system dynamics are known or deterministic, or that the state and action space are finite, significantly limiting application scope. We address this limitation by developing a probabilistic verification framework for unknown dynamical systems which combines conformal prediction with reachability analysis. In particular, we use conformal prediction to obtain valid uncertainty intervals for the unknown dynamics at each time step, with reachability then verifying whether safety is maintained within the conformal uncertainty bounds. Next, we develop an algorithmic approach for training control policies that optimize nominal reward while also maximizing the planning horizon with sound probabilistic safety guarantees. We evaluate the proposed approach in seven safe control settings spanning four domains -- cartpole, lane following, drone control, and safe navigation -- for both affine and nonlinear safety specifications. Our experiments show that the policies we learn achieve the strongest provable safety guarantees while still maintaining high average reward.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。