通过扰动注入增强模仿学习安全性,让机器人更安全地应对突发状况。
Safety-Aware Imitation Learning via MPC-Guided Disturbance Injection
- 用基于采样的MPC生成最坏情况扰动,模拟危险场景。
- 在仿真和真实四旋翼上验证,安全性和任务表现均提升。
- 适合需要高安全性的机器人控制应用,如飞行器、移动机器人。
模仿学习为从专家示范中学习复杂机器人行为提供了有效途径。然而,学习到的策略可能因错误导致安全违规,限制其在关键安全场景中的部署。本文提出MPC-SafeGIL,一种设计阶段的安全增强方法:在专家示范过程中注入对抗性扰动,使专家暴露于更广泛的危急场景中,从而让模仿策略学会鲁棒的恢复行为。该方法采用基于采样的模型预测控制(MPC)近似最坏情况扰动,可扩展至高维及黑箱动力系统。与依赖解析模型或交互式专家的前期工作不同,MPC-SafeGIL将安全考量直接融入数据采集过程。我们在四足步态和视觉-运动导航的大量仿真以及真实四旋翼实验中验证了该方法,结果表明其在安全性与任务性能上均有显著提升。
原文摘要 · Abstract (English)
Imitation Learning has provided a promising approach to learning complex robot behaviors from expert demonstrations. However, learned policies can make errors that lead to safety violations, which limits their deployment in safety-critical applications. We propose MPC-SafeGIL, a design-time approach that enhances the safety of imitation learning by injecting adversarial disturbances during expert demonstrations. This exposes the expert to a broader range of safety-critical scenarios and allows the imitation policy to learn robust recovery behaviors. Our method uses sampling-based Model Predictive Control (MPC) to approximate worst-case disturbances, making it scalable to high-dimensional and black-box dynamical systems. In contrast to prior work that relies on analytical models or interactive experts, MPC-SafeGIL integrates safety considerations directly into data collection. We validate our approach through extensive simulations including quadruped locomotion and visuomotor navigation and real-world experiments on a quadrotor, demonstrating improvements in both safety and task performance. See our website here: https://leqiu2003.github.io/MPCSafeGIL/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。