arXiv:2506.16685cs.ROcs.LG2025-06NeurIPS被引 47

用人类轻柔修正提升机器人抓取复杂操作成功率

Compliant Residual DAgger: Improving Real-World Contact-Rich Manipulation with Human Corrections

  • 通过柔性干预接口让人类在不打断机器人的前提下提供精准动作修正
  • 仅用少量修正数据使基础策略成功率提升64%,在四个任务中表现最优
  • 适合需要高精度接触操作的工业机器人学习场景

我们解决现实世界中高接触频率操作任务中数据聚合(DAgger)的关键挑战:如何获取有信息量的人类修正数据,以及如何有效利用这些数据更新策略。提出柔性残差DAgger(CR-DAgger),包含两个新组件:1)柔性干预接口,利用顺应性控制使人类能在不中断机器人执行的情况下提供温和、精确的动作增量修正;2)柔性残差策略,从人类修正中学习的同时融合力反馈与力控制。实验表明,该系统在四个具有挑战性的任务(书本翻转、皮带装配、电缆布线、齿轮插入)中,仅用少量修正数据即显著提升性能,使基础策略成功率平均提高64%,优于从头训练和微调方法。通过大量真实世界实验,为实际机器人学习中的高效DAgger实现提供了实用指导。视频演示见:https://compliant-residual-dagger.github.io

原文摘要 · Abstract (English)

We address key challenges in Dataset Aggregation (DAgger) for real-world contact-rich manipulation: how to collect informative human correction data and how to effectively update policies with this new data. We introduce Compliant Residual DAgger (CR-DAgger), which contains two novel components: 1) a Compliant Intervention Interface that leverages compliance control, allowing humans to provide gentle, accurate delta action corrections without interrupting the ongoing robot policy execution; and 2) a Compliant Residual Policy formulation that learns from human corrections while incorporating force feedback and force control. Our system significantly enhances performance on precise contact-rich manipulation tasks using minimal correction data, improving base policy success rates by 64% on four challenging tasks (book flipping, belt assembly, cable routing, and gear insertion) while outperforming both retraining-from-scratch and finetuning approaches. Through extensive real-world experiments, we provide practical guidance for implementing effective DAgger in real-world robot learning tasks. Result videos are available at: https://compliant-residual-dagger.github.io

机器人操作人类修正力控学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。