用人类第一视角视频预训练,效果比真实机器人数据更好。
HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining

- 用筛选与标注管道处理人类视角视频,替代真实机器人数据
- 相同数据量下,验证损失降低24%,任务成功率提升超90%
- 适合想低成本获取多样化世界表征的研究者
具身基础模型虽有望像大语言模型一样通过数据规模提升性能,却受限于数据瓶颈。当前主流仍依赖高成本、低多样性的远程操控机器人轨迹,而人类第一视角视频因其低成本、高多样性成为潜在替代。本文系统比较了该类视频与真实机器人数据在具身模型预训练中的表现,结果出人意料:经精心设计的过滤与标注流程处理后,人类视角视频不仅可作为有效替代,更带来更优性能——在相同预训练数据量下,模型在真实机器人动作预测上验证损失降低24%,在分布内和分布外任务执行的成功率分别提升52.5%和90%。研究验证了一种新范式:先用人类视角视频学习多样化世界表征,再以少量标注机器人数据进行动作空间对齐。本研究推动了对人类视角数据的探索,并为机器人数据采集前的质量评估提供指导。
原文摘要 · Abstract (English)
Embodied foundation models are expected to benefit from data scaling like large language models, but face a much tighter data bottleneck. Teleoperated real-robot trajectories remain the dominant pretraining source due to their precise action supervision and embodiment alignment, yet their scalability is limited by high collection cost, acquisition difficulty, and low behavioral and environmental diversity. These limitations have sparked interest in egocentric human video as a scalable, substantially lower-cost, and more diverse alternative for embodied model pretraining. However, its effectiveness compared to teleoperated real-robot data remains underexplored. To address this question, we conduct a systematic study comparing egocentric human video and teleoperated real-robot trajectories as pretraining data sources for embodied foundation models, under fixed post-training and validation protocols. Surprisingly, we find that egocentric data, when processed through a carefully designed filtering and labeling pipeline, is not merely a viable substitute for model pretraining but can lead to superior performance. With the same amount of pretraining data, models pretrained on egocentric data achieve a 24% lower validation loss on real-robot action prediction, as well as 52.5% and 90% higher success rates on in-distribution and out-of-distribution real-robot task execution, respectively. This finding verifies a scalable paradigm for embodied foundation models: pretrain on egocentric human video to learn diverse world representations, then adapt with a small amount of labeled real-robot data for action-space alignment. We hope this study encourages broader exploration of egocentric data and offers guidance for data quality assessment before costly robot data collection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。