将人类第一视角视频转为机器人训练数据,提升模型泛化能力
Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data

- 通过动作迁移与视觉合成,将人类视频转化为机器人可学数据
- 生成18,561小时跨15种机器人形态的训练数据,规模全球最大
- 适合研究视觉-语言-动作模型泛化与机器人数据合成的学者
学习通用机器人操作策略需要大规模、多样化的示范数据。第一人称人类操作视频蕴含丰富的场景与任务多样性,已有研究表明将此类视频重定向并渲染为机器人格式数据,可在小规模下获得有效任务策略。然而,该方法在大规模下是否对视觉-语言-动作模型具备预训练价值尚未被探索。本文提出「Ego2Robot」,一个可扩展的流水线,通过动作重定向、机器人臂视觉合成与多层级质量筛选,将第一人称人类操作视频转化为机器人训练数据。该方法支持经筛选的数据集与真实环境视频,共生成18,561小时机器人训练数据,覆盖15种机器人形态,是目前最大的从人类到机器人数据转换的公开数据集。为评估泛化能力,我们扩展RoboTwin2.0,引入解耦扰动轴,涵盖视觉外观、场景布局、本体形态和任务语义。实验表明,在Ego2Robot合成数据与真实机器人数据上联合预训练,可持续提升多种扰动类型下的分布外泛化性能,并在真实机器人部署中得到验证。
原文摘要 · Abstract (English)
Learning generalizable robot manipulation policies requires large-scale and diverse demonstration data. Egocentric human manipulation videos offer rich scene and task diversity, and prior work has shown that retargeting and rendering such videos into robot-format data can yield effective per-task policies at small scale. However, whether this approach can provide pretraining benefits for vision-language-action models at scale remains unexplored. We present \textbf{Ego2Robot}, a scalable pipeline that converts egocentric human manipulation videos into robot training data through action retargeting, robot-arm visual synthesis, and multi-level quality curation. Ego2Robot supports both curated datasets and in-the-wild videos, producing 18,561 hours of robot training data spanning 15 robot morphologies, making it the largest ego-to-robot dataset to date. To evaluate generalization, we extend RoboTwin2.0 with disentangled perturbation axes covering visual appearance, scene layout, embodiment morphology, and task semantics. Experiments show that joint pretraining on Ego2Robot-synthesized and robot data consistently improves out-of-distribution generalization across multiple perturbation types, with benefits validated on real-robot deployment. Project page: https://www-ye.github.io/ego2robot_blog/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。