latent视频预测模型更擅长构建鲁棒的世界模型
Latent Video Prediction Learns Better World Models

- 用隐空间预测替代像素重建,提升模型对扰动的适应能力
- 在遮挡和噪声下仍能保持类别结构与时间方向感知
- 适合需要可靠物理理解的视频建模任务
自监督视频模型正被视作世界模型,但其评估仍局限于干净基准上的单一准确率。本文首次系统研究这一空白,对比了四种容量相当的前沿视频基础模型(V-JEPA 2.1、V-JEPA 2、VideoPrism、VideoMAEv2),在五项与部署相关的关键鲁棒性维度上进行评估:特征可区分性、噪声鲁棒性、细粒度辨别能力、遮挡鲁棒性以及对时间方向的敏感性。结果表明,隐空间预测模型在所有五项指标上均表现出一致且独特的优势:在像素干扰下退化更平缓;遮挡时保留可用类别结构而非仅几何稳定;无需像素重建即可捕捉细微物理接触线索;并唯一编码时间箭头。这些优势在任务适配后依然存在:冻结的V-JEPA 2主干搭配轻量注意力探针,在噪声与遮挡鲁棒性上超越全微调的VideoMAE和监督式TimeSformer。实验为隐空间预测在鲁棒世界建模中的有效性提供了坚实证据。
原文摘要 · Abstract (English)
Self-supervised video models are increasingly framed as world models, yet their evaluation remains largely confined to a single top-1 accuracy score on clean benchmarks. This leaves a major gap in comprehending their potential as world models. We present the first systematic study addressing this gap, analyzing four matched-capacity frontier video foundation models, V-JEPA 2.1, V-JEPA 2, VideoPrism, and VideoMAEv2, across five robustness axes relevant to their deployment as video world models: feature discriminability, corruption robustness, fine-grained discrimination, occlusion robustness, and sensitivity to temporal direction. Our evaluations establish that across all five axes, latent-prediction models form a distinct and consistent profile. They degrade more gracefully under pixel corruption, preserve usable class structure rather than mere geometric stability under occlusion, capture fine-grained physical contact cues without reconstructing pixels, and uniquely encode the arrow of time. These advantages can even survive task adaptation: a frozen V-JEPA 2 backbone with a lightweight attentive probe outperforms a fully fine-tuned VideoMAE and a supervised TimeSformer on corruption and occlusion robustness. Our extensive results offer concrete new evidence in favor of latent prediction for robust world modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。