arXiv:2603.19235cs.CVcs.RO2026-03中稿 · ECCV被引 11

用视频生成模型的隐式空间先验,让大模型学会理解三维场景。

Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding

论文配图:Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding
图 1 · 摘自论文原文
  • 利用视频生成模型内部隐含的3D结构与物理规律知识
  • 在多个场景理解任务中超越现有方法,提升几何推理能力
  • 无需额外3D标注,适合需要空间认知的AI应用

尽管多模态大语言模型具备出色的语义能力,但在细粒度几何推理和物理动态理解方面常表现出空间盲区。现有方法通常依赖显式的3D模态或复杂的几何框架,受限于数据稀缺与泛化挑战。本文提出范式转变:利用大规模视频生成模型中的隐式空间先验。我们认为,为生成时序连贯视频,这些模型自然习得稳健的3D结构先验与物理法则。我们提出VEGA-3D(Video Extracted Generative Awareness)——一种即插即用框架,将预训练视频扩散模型重用于潜在世界模拟器。通过从中间噪声层级提取时空特征,并结合语义表示,采用令牌级自适应门控融合机制,我们在不依赖显式3D监督的前提下,向大模型注入密集几何线索。大量实验在3D场景理解、空间推理及具身操作基准上验证,本方法显著优于当前最优基线,证明生成先验可作为物理世界理解的可扩展基础。代码公开于 https://github.com/H-EmbodVis/VEGA-3D。

原文摘要 · Abstract (English)

While Multimodal Large Language Models demonstrate impressive semantic capabilities, they often suffer from spatial blindness, struggling with fine-grained geometric reasoning and physical dynamics. Existing solutions typically rely on explicit 3D modalities or complex geometric scaffolding, which are limited by data scarcity and generalization challenges. In this work, we propose a paradigm shift by leveraging the implicit spatial prior within large-scale video generation models. We posit that to synthesize temporally coherent videos, these models inherently learn robust 3D structural priors and physical laws. We introduce VEGA-3D (Video Extracted Generative Awareness), a plug-and-play framework that repurposes a pre-trained video diffusion model as a Latent World Simulator. By extracting spatiotemporal features from intermediate noise levels and integrating them with semantic representations via a token-level adaptive gated fusion mechanism, we enrich MLLMs with dense geometric cues without explicit 3D supervision. Extensive experiments across 3D scene understanding, spatial reasoning, and embodied manipulation benchmarks demonstrate that our method outperforms state-of-the-art baselines, validating that generative priors provide a scalable foundation for physical-world understanding. Code is publicly available at https://github.com/H-EmbodVis/VEGA-3D.

三维理解生成先验视频生成具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。