arXiv:2509.09769cs.RO2025-09被引 17

用真人玩耍视频教会人形机器人快速学会新操作任务

MimicDroid: In-Context Learning for Humanoid Robot Manipulation from Human Play Videos

  • 从自由互动视频中提取相似动作对,训练模型实现上下文学习
  • 真实世界测试中成功率提升近一倍,优于现有最先进方法
  • 适合研究人形机器人少样本学习与视觉泛化能力的团队

我们旨在让双足机器人仅凭少量视频示例即可高效完成新操作任务。上下文学习(ICL)因其测试时数据高效性和快速适应性,是实现该目标的有前景框架。然而,当前ICL方法依赖耗时的人类遥操作数据训练,限制了可扩展性。本文提出使用人类玩耍视频——连续、未标注的自由环境交互视频——作为可扩展且多样化的训练数据源。我们构建MimicDroid,仅用人类玩耍视频即可实现人形机器人在测试时的上下文学习。该模型通过提取具有相似操作行为的轨迹对,并训练策略以另一轨迹条件预测动作,从而获得在测试时适应新物体和环境的能力。为弥合具身差异,MimicDroid首先利用运动学相似性将从RGB视频估计的人类手腕姿态重定向至人形机器人;同时在训练中引入随机区域遮蔽,降低对人类特有线索的过拟合,提升对视觉差异的鲁棒性。为评估人形机器人少样本学习能力,我们引入一个开源仿真基准,逐步增加泛化难度。MimicDroid在真实世界中表现优于现有最先进方法,成功率接近翻倍。

原文摘要 · Abstract (English)

We aim to enable humanoid robots to efficiently solve new manipulation tasks from a few video examples. In-context learning (ICL) is a promising framework for achieving this goal due to its test-time data efficiency and rapid adaptability. However, current ICL methods rely on labor-intensive teleoperated data for training, which restricts scalability. We propose using human play videos -- continuous, unlabeled videos of people interacting freely with their environment -- as a scalable and diverse training data source. We introduce MimicDroid, which enables humanoids to perform ICL using human play videos as the only training data. MimicDroid extracts trajectory pairs with similar manipulation behaviors and trains the policy to predict the actions of one trajectory conditioned on the other. Through this process, the model acquired ICL capabilities for adapting to novel objects and environments at test time. To bridge the embodiment gap, MimicDroid first retargets human wrist poses estimated from RGB videos to the humanoid, leveraging kinematic similarity. It also applies random patch masking during training to reduce overfitting to human-specific cues and improve robustness to visual differences. To evaluate few-shot learning for humanoids, we introduce an open-source simulation benchmark with increasing levels of generalization difficulty. MimicDroid outperformed state-of-the-art methods and achieved nearly twofold higher success rates in the real world. Additional materials can be found on: ut-austin-rpl.github.io/MimicDroid

人形机器人上下文学习视频理解少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。