用真人日常视频训练机器人操控模型,零样本即能执行新任务。
Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
- 用自动化方法从真人第一视角视频中提取动作片段与语言描述。
- 构建含100万条任务、2600万帧的大型手部操作数据集。
- 仅需少量真实机器人数据微调,就能在新物体上实现高成功率。
本文提出一种新方法,利用大量未经脚本化的真实人类手部活动第一视角视频预训练机器人视觉-语言-动作(VLA)模型。将人类手部视为灵巧机器人末端执行器,通过全自动的人类活动分析系统,将无标注的“野生”视频转化为与现有机器人VLA训练数据对齐的数据格式,包含原子级动作片段、语言描述及逐帧3D手部与相机运动。我们处理了海量第一视角视频,构建了一个包含100万条任务实例和2600万帧的双手部VLA训练数据集,覆盖丰富多样的物体、概念、灵巧操作任务及真实环境变化,远超现有机器人数据集的覆盖范围。设计并预训练了一个灵巧手部VLA模型,该模型在完全未见过的真实观测下表现出强大零样本能力。此外,在少量真实机器人动作数据上微调后,任务成功率显著提升,并展现出对新物体的良好泛化能力。还验证了模型性能随预训练数据规模增长而持续提升的可扩展性。我们认为该工作为可扩展的VLA预训练奠定了坚实基础,推动机器人向真正通用的具身智能迈进。
原文摘要 · Abstract (English)
This paper presents a novel approach for pretraining robotic manipulation Vision-Language-Action (VLA) models using a large corpus of unscripted real-life video recordings of human hand activities. Treating human hand as dexterous robot end-effector, we show that "in-the-wild" egocentric human videos without any annotations can be transformed into data formats fully aligned with existing robotic V-L-A training data in terms of task granularity and labels. This is achieved by the development of a fully-automated holistic human activity analysis approach for arbitrary human hand videos. This approach can generate atomic-level hand activity segments and their language descriptions, each accompanied with framewise 3D hand motion and camera motion. We process a large volume of egocentric videos and create a hand-VLA training dataset containing 1M episodes and 26M frames. This training data covers a wide range of objects and concepts, dexterous manipulation tasks, and environment variations in real life, vastly exceeding the coverage of existing robot data. We design a dexterous hand VLA model architecture and pretrain the model on this dataset. The model exhibits strong zero-shot capabilities on completely unseen real-world observations. Additionally, fine-tuning it on a small amount of real robot action data significantly improves task success rates and generalization to novel objects in real robotic experiments. We also demonstrate the appealing scaling behavior of the model's task performance with respect to pretraining data scale. We believe this work lays a solid foundation for scalable VLA pretraining, advancing robots toward truly generalizable embodied intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。