arXiv:2508.09976cs.RO2025-08被引 54

用真人视频生成机器人可学的演示,提升泛化能力

Masquerade: Learning from In-the-wild Human Videos using Data-Editing

  • 将真人第一视角视频转为带机器人的模拟演示
  • 仅用50个机器人示范+67.5万帧编辑视频训练,效果提升5-6倍
  • 适合想用公开视频数据训练机器人策略的研究者

机器人操作研究仍面临严重数据稀缺问题:即使最大规模的机器人数据集也远小于推动语言和视觉领域突破的海量数据。我们提出Masquerade方法,通过编辑真实世界中第一视角的人类视频来弥合人类与机器人之间的视觉具身差距,并据此学习机器人策略。该流程将每段人类视频转化为机器人演示:(i) 估计3维手部姿态,(ii) 填充人类手臂区域,(iii) 叠加渲染的双臂机器人模型并跟踪恢复出的末端执行器轨迹。在67.5万帧此类编辑视频上预训练视觉编码器以预测未来2维机器人关键点,并在微调扩散策略头时继续使用该辅助损失,仅需每个任务50个机器人示范,即可获得显著优于以往工作的泛化策略。在三个长程、双臂厨房任务中,于三个未见场景评估,性能比基线高出5-6倍。消融实验表明,机器人叠加与联合训练均不可或缺,性能随编辑人类视频量呈对数增长。结果证明,显式缩小视觉具身差距,可释放大量现成可用的人类视频数据,用于提升机器人策略。

原文摘要 · Abstract (English)

Robot manipulation research still suffers from significant data scarcity: even the largest robot datasets are orders of magnitude smaller and less diverse than those that fueled recent breakthroughs in language and vision. We introduce Masquerade, a method that edits in-the-wild egocentric human videos to bridge the visual embodiment gap between humans and robots and then learns a robot policy with these edited videos. Our pipeline turns each human video into robotized demonstrations by (i) estimating 3-D hand poses, (ii) inpainting the human arms, and (iii) overlaying a rendered bimanual robot that tracks the recovered end-effector trajectories. Pre-training a visual encoder to predict future 2-D robot keypoints on 675K frames of these edited clips, and continuing that auxiliary loss while fine-tuning a diffusion policy head on only 50 robot demonstrations per task, yields policies that generalize significantly better than prior work. On three long-horizon, bimanual kitchen tasks evaluated in three unseen scenes each, Masquerade outperforms baselines by 5-6x. Ablations show that both the robot overlay and co-training are indispensable, and performance scales logarithmically with the amount of edited human video. These results demonstrate that explicitly closing the visual embodiment gap unlocks a vast, readily available source of data from human videos that can be used to improve robot policies.

机器人策略视频编辑具身学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。