提出安全保证机制,让机器人在不冒险的前提下最大化任务成功率。
To Do or Not to Do: Ensuring the Safety of Visuomotor Policies Learned from Demonstrations

- 用视图合成识别状态空间中的安全区域,结合不变集理论确保执行可靠性。
- 实验表明,即使有运行时扰动,也能在安全区域内实现最大任务成功率。
- 附带的恢复策略可提升性能,缓解安全与成功率的矛盾,适合实际部署场景。
任务成功一直是模仿学习(IL)中衡量策略性能的主要指标,但这一特性严重限制了其在实地机器人应用中的普及,因为安全性同样至关重要。理想情况下,若无法保障安全,机器人不应执行可能失败的策略。尽管经典控制领域已深入研究安全与性能的权衡,但该问题在模仿学习中仍鲜有探讨,且缺乏通用的安全定义。现有理论也难以扩展至实际机器人。本文提出「执行保证」——一种不依赖具体策略的安全度量方法,可在状态空间特定区域内确保视觉运动策略的最大任务成功率,即便存在微小扰动。通过利用最近的视图合成技术定位这些安全区域,并基于集合不变性中的纳古莫条件(Nagumo's sub-tangentiality condition),实现并验证该保证。在Franka机械臂的仿真与真实世界实验中,证明所提分析可使多种模仿学习策略在保证安全的前提下实现最大成功率。此外,还展示了由该分析衍生的恢复策略能有效提升性能,从而缓解安全与性能间的权衡。
原文摘要 · Abstract (English)
Task success has historically been the primary measure of policy performance in imitation learning (IL) research. This characteristics strictly limits the ubiquitous applications of IL algorithms in field robotics where safety assurance, in addition to task-success, is of paramount importance. It is often desirable for an IL-powered robot in the field not to roll out a policy, and hence score a poor performance, if the safety is not guaranteed. Although this trade-off between safety and performance is well investigated in classical control literature, policy safety is a heavily underexplored domain in IL research. There is no universal definition of safety in IL. To make things worst, many existing theoretical works on safety is notoriously difficult to extend to IL-powered robots in the field. This paper offers important insights on the safety and performance of IL policies. We propose execution guarantee, a policy-agnostic safety measure that guarantees the maximum task success for a visuomotor IL policy, despite minor run-time changes, from within a specific region in the state space. We leverage recent advances in view synthesis to identify such regions in the state space for an IL policy and explore a fundamental result on set invariance - namely, Nagumo's sub-tangentiality condition - to prove and operationalize execution guarantee from inside that region. Experiments with a Franka robot, both in simulation and real world, demonstrate how the proposed safety analysis allows various IL policies to achieve maximum task success with guarantee. We also demonstrate some interesting results on how a recovery policy - a by-product of the proposed safety analysis - can help to increase the policy performance and thereby mitigating the safety-performance tradeoff in IL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。