arXiv:2606.09646cs.CVcs.AI2026-06被引 1

探究视频大模型是否掌握直觉物理,发现其表征随深度增强且依赖训练方式。

Do Video Foundation Models Understand Intuitive Physics? A Layerwise Probing Analysis

论文配图:Do Video Foundation Models Understand Intuitive Physics? A Layerwise Probing Analysis
图 1 · 摘自论文原文
  • 用分层探测法分析冻结特征,比较三类模型对物理规律的编码能力。
  • 中间到深层特征中物理信息最易提取,破坏帧序会显著降低性能。
  • 基于动态建模的探测器表现最优,适合研究视觉认知与生成模型机制。

我们研究预训练视频基础模型在其冻结表征中是否编码了直觉物理信息,以及该信息如何随模型家族、层级和探测类型变化。在IntPhys2和最小视频对(MVP)数据集上,使用冻结特征探测法对比了预测联合嵌入模型(V-JEPA)、掩码重建模型(VideoMAE)和基于扩散的视频生成器(LTX-Video)。V-JEPA在所有基准测试中表现最强,尤其在建模时间动态的探测器下;VideoMAE保持竞争力,而LTX-Video虽较弱但仍能恢复非平凡信号。分层分析显示,物理相关信息在早期层最弱,于中后期层次最易获取;时间控制实验表明,破坏帧序会显著降低性能,尤其在MVP上。结果共同表明,直觉物理知识可稳定出现在预训练视频表征中,但其可访问性高度依赖预训练范式、表征深度及读出机制。

原文摘要 · Abstract (English)

We study whether pretrained video foundation models encode intuitive-physics information in their frozen representations, and how this information varies across model families, layers, and probe types. Using frozen-feature probing on IntPhys2 and Minimal Video Pairs (MVP), we compare predictive joint-embedding models (V-JEPA), masked reconstruction models (VideoMAE), and a diffusion-based video generator (LTX-Video). V-JEPA achieves the strongest overall results across benchmarks, especially with probes that model temporal dynamics, while VideoMAE remains competitive and LTX-Video recovers weaker but non-trivial signal. Layerwise analyses show that physics-relevant information is weakest in early layers and becomes most accessible at intermediate-to-late depth, and temporal controls show that disrupting frame order substantially reduces performance, especially on MVP. Together, these results suggest that intuitive-physics knowledge emerges reliably in pretrained video representations, but its accessibility depends strongly on pretraining paradigm, representational depth, and readout mechanism.

视频理解直觉物理表征分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。