让机器人从失败中学习,提升导航安全性。
Learning from Demonstration with Failure Awareness for Safe Robot Navigation

- 分离成功与失败数据,失败数据用于识别危险区域
- 碰撞率显著降低,任务成功率保持不变
- 适用于不同机器人平台,真实环境效果好
示范学习广泛应用于机器人导航,但存在根本局限:示范数据以成功行为为主,难以覆盖不安全状态。这导致机器人在超出示范分布的场景中表现不佳。碰撞等失败经历蕴含关键的危险区域信息,却常被忽视。其难点在于失败数据无法直接指导动作模仿,若盲目纳入策略学习会损害性能。为此,本文提出一种失败感知学习框架,显式分离成功与失败数据的作用:失败经验用于构建危险区域的价值估计,而策略学习仅基于成功示范。该设计使失败数据有效利用而不干扰策略行为。我们在离线强化学习设置下实现此框架,并在仿真与真实环境评估。结果表明,本方法持续降低碰撞率,同时维持任务成功率,在不同环境与机器人平台上均展现良好泛化能力。
原文摘要 · Abstract (English)
Learning from demonstration is widely used for robot navigation, yet it suffers from a fundamental limitation: demonstrations consist predominantly of successful behaviors and provide limited coverage of unsafe states. This limitation leads to poor safety when the robot encounters scenarios beyond the demonstration distribution. Failure experiences, such as collisions, contain essential information about unsafe regions, but remain underutilized. The key difficulty lies in the fact that failure data do not provide valid guidance for action imitation, and their naive incorporation into policy learning often degrades performance. We address this challenge by proposing a failure-aware learning framework that explicitly decouples the roles of success and failure data. In this framework, failure experiences are used to shape value estimation in hazardous regions, while policy learning is restricted to successful demonstrations. This separation enables the effective use of failure data without corrupting policy behavior. We implement this design within an offline reinforcement learning (RL) setting and evaluate it in both simulation and real-world environments. The results show that our framework consistently reduces collision rates while preserving the task success rate, and demonstrate strong generalization across different environments and robot platforms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。