arXiv:2603.14482cs.CV2026-03被引 67

V-JEPA 2.1通过多组件设计,实现视频图像的高质量密集视觉表征。

V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning

  • 采用掩码预测损失与分层自监督,强化时空定位能力
  • 在多个基准上达到新纪录,如Ego4D mAP达7.71,机器人抓取成功率提升20点
  • 适合需要精准时空理解的视觉任务,如机器人交互与动作预测

我们提出V-JEPA 2.1,一种自监督模型家族,可学习图像与视频的高精度密集视觉表征,同时保持强全局场景理解能力。该方法融合四大核心组件:第一,基于掩码的密集预测损失,使可见与被遮挡片段均贡献训练信号,增强空间与时间定位;第二,分层自监督机制,在多个中间编码层应用自监督目标,提升表征质量;第三,多模态分词器支持图像与视频统一训练;第四,模型容量与训练数据的有效扩展。上述设计共同生成具有空间结构、语义连贯且时序一致的表征。实验表明,V-JEPA 2.1在多个挑战性基准上表现优异:在Ego4D上短时物体交互预测达到7.71 mAP,EPIC-KITCHENS上高层动作预测召回率Recall@5为40.8,相比V-JEPA-2 AC在真实机器人抓取任务中成功率提升20个百分点。此外,在机器人导航(TartanDrive ATE 5.687)、深度估计(NYUv2 RMSE 0.307,线性探测)和全局识别(Something-Something-V2 准确率77.7)上也表现强劲。结果表明,V-JEPA 2.1显著推进了密集视觉理解与世界建模的前沿水平。

原文摘要 · Abstract (English)

We present V-JEPA 2.1, a family of self-supervised models that learn dense, high-quality visual representations for both images and videos while retaining strong global scene understanding. The approach combines four key components. First, a dense predictive loss uses a masking-based objective in which both visible and masked tokens contribute to the training signal, encouraging explicit spatial and temporal grounding. Second, deep self-supervision applies the self-supervised objective hierarchically across multiple intermediate encoder layers to improve representation quality. Third, multi-modal tokenizers enable unified training across images and videos. Finally, the model benefits from effective scaling in both model capacity and training data. Together, these design choices produce representations that are spatially structured, semantically coherent, and temporally consistent. Empirically, V-JEPA 2.1 achieves state-of-the-art performance on several challenging benchmarks, including 7.71 mAP on Ego4D for short-term object-interaction anticipation and 40.8 Recall@5 on EPIC-KITCHENS for high-level action anticipation, as well as a 20-point improvement in real-robot grasping success rate over V-JEPA-2 AC. The model also demonstrates strong performance in robotic navigation (5.687 ATE on TartanDrive), depth estimation (0.307 RMSE on NYUv2 with a linear probe), and global recognition (77.7 on Something-Something-V2). These results show that V-JEPA 2.1 significantly advances the state of the art in dense visual understanding and world modeling.

视频自监督密集表征机器人感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。