探究视频模型是否自动掌握3D知识,发现顶尖生成模型有超越专业3D模型的3D理解能力。
How Much 3D Do Video Foundation Models Encode?
- 设计通用框架,通过浅层读出估计特征中的3D属性来量化模型3D意识。
- 顶尖视频生成模型虽未接触3D数据,却展现出强于专门3D模型的3D理解能力。
- 为构建可扩展3D模型提供关键洞察,适合关注多模态理解与3D感知的研究者。
视频是3D世界在2D平面上的连续投影。在大量视频数据上训练后,全局3D理解是否会自然涌现?我们通过量化现有视频基础模型(VidFMs)在海量视频数据预训练后的3D理解能力,研究这一问题。提出首个模型无关框架,通过浅层读出从模型特征中估计多种3D属性,以衡量不同VidFMs的3D意识。研究在多个维度揭示了重要发现:最先进的视频生成模型尽管未接受任何3D数据训练,却表现出对3D物体与场景的强理解能力,甚至超越专门针对3D任务训练的大规模专家模型。这些发现结合主要VidFMs的3D基准测试结果,为构建可扩展3D模型提供了宝贵洞见。
原文摘要 · Abstract (English)
Videos are continuous 2D projections of 3D worlds. After training on large video data, will global 3D understanding naturally emerge? We study this by quantifying the 3D understanding of existing Video Foundation Models (VidFMs) pretrained on vast video data. We propose the first model-agnostic framework that measures the 3D awareness of various VidFMs by estimating multiple 3D properties from their features via shallow read-outs. Our study presents meaningful findings regarding the 3D awareness of VidFMs on multiple axes. In particular, we show that state-of-the-art video generation models exhibit a strong understanding of 3D objects and scenes, despite not being trained on any 3D data. Such understanding can even surpass that of large expert models specifically trained for 3D tasks. Our findings, together with the 3D benchmarking of major VidFMs, provide valuable observations for building scalable 3D models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。