arXiv:2509.07953cs.ROcs.LG2025-09被引 40

用人类干预数据训练机器人学会重试与修正,提升长时任务成功率。

RaC: Robot Learning for Long-Horizon Tasks by Scaling Recovery and Correction

  • 在模仿学习后引入人类介入的回放阶段,学习恢复与修正行为
  • 仅需1/10的数据收集时间,便在多个真实任务中超越现有最优表现
  • 支持测试时扩展:恢复动作越多,性能越强,适合复杂长时任务

当前机器人模仿学习虽在大量人类示范数据上训练出强大策略,但在接触密集、柔性物体及长时任务中仍远未达到完美执行。这源于依赖人工遥操作收集的‘专家’数据效率低下。为此,我们提出RaC——在模仿学习预训练后,通过人机协同回放进行微调。在策略执行过程中,当失败临近时,人类操作者会先将机器人回退至熟悉状态,再提供修正动作完成子任务。基于此类数据训练,机器人习得重试与适应能力,显著提升长时任务中的效率与鲁棒性。在三个真实双臂任务(衬衫悬挂、密封容器盖、外卖盒打包)和一个模拟装配任务中,RaC仅用1/10的数据收集时间与样本量,即超越先前最优方法。此外,我们证明了测试时可扩展性:训练后的RaC策略性能随恢复动作数量线性增长。

原文摘要 · Abstract (English)

Modern paradigms for robot imitation train expressive policy architectures on large amounts of human demonstration data. Yet performance on contact-rich, deformable-object, and long-horizon tasks plateau far below perfect execution, even with thousands of expert demonstrations. This is due to the inefficiency of existing ``expert'' data collection procedures based on human teleoperation. To address this issue, we introduce RaC, a new phase of training on human-in-the-loop rollouts after imitation learning pre-training. In RaC, we fine-tune a robotic policy on human intervention trajectories that illustrate recovery and correction behaviors. Specifically, during a policy rollout, human operators intervene when failure appears imminent, first rewinding the robot back to a familiar, in-distribution state and then providing a corrective segment that completes the current sub-task. Training on this data composition expands the robotic skill repertoire to include retry and adaptation behaviors, which we show are crucial for boosting both efficiency and robustness on long-horizon tasks. Across three real-world bimanual control tasks: shirt hanging, airtight container lid sealing, takeout box packing, and a simulated assembly task, RaC outperforms the prior state-of-the-art using 10$\times$ less data collection time and samples. We also show that RaC enables test-time scaling: the performance of the trained RaC policy scales linearly in the number of recovery maneuvers it exhibits. Videos of the learned policy are available at https://rac-scaling-robot.github.io/.

机器人学习长时任务人机协作强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。