训练中实时验证,让控制策略可安全运行
Learning Verifiable Control Policies Using Relaxed Verification
- 用可微可达性分析,在训练时嵌入验证机制
- 在四旋翼和单轮车模型上满足避障与不变性要求
- 适合对安全性有严苛要求的自动驾驶等场景
为学习型控制系统提供安全保证,现有方法在训练结束后进行形式化验证。但若策略不满足规范,或验证算法存在保守性,则难以建立安全保障。本文提出在训练过程中持续进行验证,最终目标是获得可在运行时通过轻量级、宽松的验证算法评估特性的控制策略。方法基于可微可达性分析,将新组件融入损失函数。在四旋翼模型和单轮车模型上的数值实验表明,该方法能生成满足期望可达-避让及不变性规范的控制策略。
原文摘要 · Abstract (English)
To provide safety guarantees for learning-based control systems, recent work has developed formal verification methods to apply after training ends. However, if the trained policy does not meet the specifications, or there is conservatism in the verification algorithm, establishing these guarantees may not be possible. Instead, this work proposes to perform verification throughout training to ultimately aim for policies whose properties can be evaluated throughout runtime with lightweight, relaxed verification algorithms. The approach is to use differentiable reachability analysis and incorporate new components into the loss function. Numerical experiments on a quadrotor model and unicycle model highlight the ability of this approach to lead to learned control policies that satisfy desired reach-avoid and invariance specifications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。