人机协同框架提升少样本机器人学习效率
Human-Robot Copilot for Data-Efficient Imitation Learning
- 通过间歇性人类干预优化机器人策略,减少数据依赖
- 相同演示轨迹下性能优于现有方法,避免状态偏离
- 适配多种机械臂,兼顾操作精细度与通用性
通过遥操作收集人类示范是教授机器人特定任务技能的常见方法。然而,当示范数量有限时,策略易因累积误差或环境随机性进入分布外(OOD)状态。现有交互式模仿学习或人机协同方法多遵循人控DAgger范式,通过执行过程中的选择性人类干预来扩充示范。但这些方法难以兼顾灵巧性与泛化性:要么仅适用于特定运动结构且提供细粒度修正,要么牺牲精确控制以换取泛化能力。为此,我们提出人机协同(Human-Robot Copilot)框架,可在保持对工业与研究用机械臂广泛兼容的同时,利用可调节的缩放因子实现灵巧遥操作。实验表明,在相同示范轨迹数量下,本框架性能更优;且修正干预仅需间歇进行,显著提升了数据采集效率与耗时成本。
原文摘要 · Abstract (English)
Collecting human demonstrations via teleoperation is a common approach for teaching robots task-specific skills. However, when only a limited number of demonstrations are available, policies are prone to entering out-of-distribution (OOD) states due to compounding errors or environmental stochasticity. Existing interactive imitation learning or human-in-the-loop methods try to address this issue by following the Human-Gated DAgger (HG-DAgger) paradigm, an approach that augments demonstrations through selective human intervention during policy execution. Nevertheless, these approaches struggle to balance dexterity and generality: they either provide fine-grained corrections but are limited to specific kinematic structures, or achieve generality at the cost of precise control. To overcome this limitation, we propose the Human-Robot Copilot framework that can leverage a scaling factor for dexterous teleoperation while maintaining compatibility with a wide range of industrial and research manipulators. Experimental results demonstrate that our framework achieves higher performance with the same number of demonstration trajectories. Moreover, since corrective interventions are required only intermittently, the overall data collection process is more efficient and less time-consuming.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。