用强化学习让机器人更聪明地吸收人类纠错,即使指令不完美也能提升表现。
ROVE: Unlocking Human Interventions for Humanoid Manipulation via Reinforcement Learning

- 构建人机协同数据收集流程,实现真实机器人干预数据采集
- 通过乐观价值估计筛选高质量行为,避免学坏习惯
- 融合跨机器人经验视频,增强对罕见失败场景的学习能力
人类干预为后训练视觉-语言-动作(VLA)模型提供关键修正信号。然而,由于全身运动学复杂和灵巧手控制困难,实现流畅的人类干预面临重大系统挑战,导致收集的干预轨迹往往质量不佳,依赖人类干预作为专家监督的方法可能学习到犹豫、低效甚至错误的行为。为此,我们提出ROVE,一种面向人形机器人VLA后训练的强化学习框架,可处理不完美的人类干预。首先,ROVE引入人机协同管道,用于收集人形机器人操作任务中的部署与干预数据;其次,采用乐观价值估计(OVE)从混合质量轨迹中优先识别高价值行为。为进一步增强价值估计鲁棒性,我们引入跨体感人类经验视频,为长尾失败与恢复模式提供丰富监督。由此生成的评判器输出具有信息量的优势信号,引导VLA策略聚焦于高价值行为,而非盲目模仿所有动作。在具有挑战性的现实世界接触密集型与精细操作任务上,ROVE显著优于基于经验学习的基线方法,并在多次回放-干预迭代中持续提升性能。
原文摘要 · Abstract (English)
Human interventions provide crucial corrective signals for post-training Vision-Language-Action (VLA) models. However, enabling seamless humanoid interventions is a formidable systems challenge due to complex whole-body kinematics and dexterous-hand control. Consequently, the collected intervention trajectories are often suboptimal, and methods that rely on human interventions as expert supervision can absorb hesitant, inefficient, or even erroneous behaviors. To address both the system and algorithmic challenges, we propose ROVE, a reinforcement learning framework for humanoid VLA post-training with imperfect human interventions. First, ROVE introduces a human-in-the-loop pipeline capable of collecting deployment and intervention data for humanoid manipulation. Second, it utilizes Optimistic Value Estimation (OVE) to prioritize high-value behaviors from mixed-quality trajectories. To further robustify value estimation, we incorporate cross-embodiment human experience videos to provide rich supervision for long-tailed failure and recovery modes. The resulting critic yields informative advantage signals, steering the VLA actor to focus on high-value behaviors rather than indiscriminately imitating all actions. On challenging real-world contact-rich and fine-grained humanoid manipulation tasks, ROVE outperforms experience-learning baselines and consistently improves across multiple rollout-intervention iterations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。