用海量视频+少量机器人数据,训练出能理解、预测、规划的通用视觉模型。
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

- 基于百万小时网络视频预训练无动作联合嵌入模型,实现跨任务理解。
- 在动作预测和问答任务中达到顶尖性能,如Epic-Kitchens-100 recall-at-5达39.7。
- 仅用62小时机器人视频微调,即可零样本部署于机械臂完成抓取任务。
现代AI面临的核心挑战是通过观察学习理解世界并做出决策。本文提出一种自监督方法,结合互联网规模视频数据与少量交互数据(机器人轨迹),构建具备理解、预测与规划能力的模型。首先,在包含超过100万小时网络视频的图像与视频数据集上预训练无动作联合嵌入预测架构V-JEPA 2,其在运动理解任务中取得77.3%的top-1准确率(Something-Something v2),在人类动作预测任务中达到39.7%的recall-at-5(Epic-Kitchens-100),超越此前专用模型。进一步将V-JEPA 2与大型语言模型对齐后,在多个视频问答任务中展现领先表现(80亿参数级:PerceptionTest 84.0,TempCompass 76.9)。最后,通过在Droid数据集的少于62小时未标注机器人视频上微调,构建潜空间动作条件世界模型V-JEPA 2-AC,零样本部署至两个实验室的Franka机械臂,实现基于图像目标的抓放规划,全程无需环境数据采集或特定任务训练与奖励设计。该工作证明,仅需网络视频与少量机器人数据,即可获得可执行物理世界规划的通用世界模型。
原文摘要 · Abstract (English)
A major challenge for modern AI is to learn to understand the world and learn to act largely by observation. This paper explores a self-supervised approach that combines internet-scale video data with a small amount of interaction data (robot trajectories), to develop models capable of understanding, predicting, and planning in the physical world. We first pre-train an action-free joint-embedding-predictive architecture, V-JEPA 2, on a video and image dataset comprising over 1 million hours of internet video. V-JEPA 2 achieves strong performance on motion understanding (77.3 top-1 accuracy on Something-Something v2) and state-of-the-art performance on human action anticipation (39.7 recall-at-5 on Epic-Kitchens-100) surpassing previous task-specific models. Additionally, after aligning V-JEPA 2 with a large language model, we demonstrate state-of-the-art performance on multiple video question-answering tasks at the 8 billion parameter scale (e.g., 84.0 on PerceptionTest, 76.9 on TempCompass). Finally, we show how self-supervised learning can be applied to robotic planning tasks by post-training a latent action-conditioned world model, V-JEPA 2-AC, using less than 62 hours of unlabeled robot videos from the Droid dataset. We deploy V-JEPA 2-AC zero-shot on Franka arms in two different labs and enable picking and placing of objects using planning with image goals. Notably, this is achieved without collecting any data from the robots in these environments, and without any task-specific training or reward. This work demonstrates how self-supervised learning from web-scale data and a small amount of robot interaction data can yield a world model capable of planning in the physical world.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。