arXiv:2409.09883cs.RO2024-09被引 2

机器人学会识别危险指令并自动推荐安全替代方案

Robots that Suggest Safe Alternatives

  • 将安全建议转化为目标空间的可控优化问题
  • 通过可达性分析构建可量化安全与存活性的价值网络
  • 在仿真中验证了推荐方案既安全又符合人类接受度

目标条件策略(如通过模仿学习获得)为人类控制机器人任务提供了便捷方式,但无法保证在分布外请求下执行安全或成功。本文提出一种安全替代方案(SALT)框架,使机器人能判断是否可安全执行用户目标,并在无法安全执行时自动建议安全替代方案。该方法受控制理论中安全过滤思想启发,将替代建议问题从动作空间转为目标空间的安全部分。离线阶段,利用可达性分析计算一个目标参数化的“可达-避障”价值网络,以量化预训练策略的安全性和存活性;在线阶段,机器人使用该价值网络作为安全过滤器,实时监测用户目标并主动提出相似但满足安全约束的替代方案。我们在室内导航和Franka Panda机械臂桌面上操作的仿真环境中测试了该框架,涵盖离散与连续目标表示。结果表明,SALT能准确预测闭环执行的成功与否,比开环不确定性估计更少悲观,且提出的替代方案始终符合人类接受标准。

原文摘要 · Abstract (English)

Goal-conditioned policies, such as those learned via imitation learning, provide an easy way for humans to influence what tasks robots accomplish. However, these robot policies are not guaranteed to execute safely or to succeed when faced with out-of-distribution requests. In this work, we enable robots to know when they can confidently execute a user's desired goal, and automatically suggest safe alternatives when they cannot. Our approach is inspired by control-theoretic safety filtering, wherein a safety filter minimally adjusts a robot's candidate action to be safe. Our key idea is to pose alternative suggestion as a safe control problem in goal space, rather than in action space. Offline, we use reachability analysis to compute a goal-parameterized reach-avoid value network which quantifies the safety and liveness of the robot's pre-trained policy. Online, our robot uses the reach-avoid value network as a safety filter, monitoring the human's given goal and actively suggesting alternatives that are similar but meet the safety specification. We demonstrate our Safe ALTernatives (SALT) framework in simulation experiments with indoor navigation and Franka Panda tabletop manipulation, and with both discrete and continuous goal representations. We find that SALT is able to learn to predict successful and failed closed-loop executions, is a less pessimistic monitor than open-loop uncertainty quantification, and proposes alternatives that consistently align with those people find acceptable.

机器人安全控制智能决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。