arXiv:2509.21986cs.ROcs.AI2025-09被引 9

用第一视角视频训练视觉语言动作模型,无需人工标注。

Developing Vision-Language-Action Model from Egocentric Videos

  • 通过EgoScaler从第一视角视频中自动提取物体6自由度运动轨迹。
  • 在模拟和真实机器人上,预训练使任务成功率提升超20%。
  • 适合想低成本构建机器人操作模型的研究者。

第一视角视频记录了人类操作物体和工具的动态过程,为学习物体操作提供了丰富的运动线索。与通常依赖专家手动遥控、成本高昂的训练方式不同,第一视角视频提供了可扩展的替代方案。然而,以往利用此类视频训练机器人策略的研究通常需要额外标注,如详细的手部姿态记录。因此,目前尚不清楚视觉语言动作模型(VLAs)能否直接从原始第一视角视频中进行训练。本文提出EgoScaler框架,无需辅助记录即可从第一视角视频中提取6自由度物体操作轨迹。我们将其应用于四个大规模第一视角视频数据集,并自动修正噪声或不完整的轨迹,从而构建一个新的大规模数据集用于VLA预训练。在模拟和真实机器人环境中,使用最先进的π₀架构进行实验,得出三个关键发现:(i) 在本数据集上预训练相比从零开始训练,任务成功率提升超过20%;(ii) 性能媲美使用真实机器人数据集的结果;(iii) 将本数据集与真实机器人数据结合可进一步提升性能。结果表明,第一视角视频是推动VLA研究的一个有前景且可扩展的资源。

原文摘要 · Abstract (English)

Egocentric videos capture how humans manipulate objects and tools, providing diverse motion cues for learning object manipulation. Unlike the costly, expert-driven manual teleoperation commonly used in training Vision-Language-Action models (VLAs), egocentric videos offer a scalable alternative. However, prior studies that leverage such videos for training robot policies typically rely on auxiliary annotations, such as detailed hand-pose recordings. Consequently, it remains unclear whether VLAs can be trained directly from raw egocentric videos. In this work, we address this challenge by leveraging EgoScaler, a framework that extracts 6DoF object manipulation trajectories from egocentric videos without requiring auxiliary recordings. We apply EgoScaler to four large-scale egocentric video datasets and automatically refine noisy or incomplete trajectories, thereby constructing a new large-scale dataset for VLA pre-training. Our experiments with a state-of-the-art $π_0$ architecture in both simulated and real-robot environments yield three key findings: (i) pre-training on our dataset improves task success rates by over 20\% compared to training from scratch, (ii) the performance is competitive with that achieved using real-robot datasets, and (iii) combining our dataset with real-robot data yields further improvements. These results demonstrate that egocentric videos constitute a promising and scalable resource for advancing VLA research.

视觉语言动作第一视角机器人学习自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。