arXiv:2603.26467cs.RO2026-03

让机器人从失败中学习,用负反馈提升模仿学习成功率。

Addressing Ambiguity in Imitation Learning through Product of Experts based Negative Feedback

  • 引入专家乘积负反馈机制,利用失败经验优化策略。
  • 在模拟与真实机器人上成功率达90%提升,优于无负反馈系统。
  • 适合处理用户演示不完美、任务模糊的现实场景。

为复杂任务编程机器人通常耗时且需专业知识。模仿学习通过人类示范训练机器人,但传统方法假设示范来自单一高技能专家。在家庭机器人等实际应用中,用户示范常不理想,且混合使用用户与预训练数据,导致任务模糊。本文提出一种负反馈系统,能利用次优示范和自身失败经验解决模糊任务。该系统在模拟与真实机器人实验中均显著优于纯正向模仿学习,在真实机器人上成功率提升50%,相比无负反馈系统提升90%。新方案还展现出更高的有效性、内存效率和时间效率,优于同类负反馈方法。

原文摘要 · Abstract (English)

Programming robots to perform complex tasks is often difficult and time consuming, requiring expert knowledge and skills in robot software and sometimes hardware. Imitation learning is a method for training robots to perform tasks by leveraging human expertise through demonstrations. Typically, the assumption is that those demonstrations are performed by a single, highly competent expert. However, in many real-world applications that use user demonstrations for tasks or incorporate both user data and pretrained data, such as home robotics including assistive robots, this is unlikely to be the case. This paper presents research towards a system which can leverage suboptimal demonstrations to solve ambiguous tasks; and particularly learn from its own failures. This is a negative-feedback system which achieves significant improvement over purely positive imitation learning for ambiguous tasks, achieving a 90% improvement in success rate against a system that does not utilise negative feedback, compared to a 50% improvement in success rate when utilised on a real robot, as well as demonstrating higher efficacy, memory efficiency and time efficiency than a comparable negative feedback scheme. The novel scheme presented in this paper is validated through simulated and real-robot experiments.

模仿学习负反馈机器人任务模糊

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。