仅用视频训练大模型,就能让AI理解三维空间关系。
Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry Priors

- 用视频生成3D几何先验,注入大模型提升空间推理能力
- 40亿参数模型在无3D数据情况下超越Gemini-1.5-Pro
- 适合想用视频做3D理解的视觉-语言研究者
以往研究将3D场景理解视为视频处理任务,通常依赖点云或重建的鸟瞰图等完整3D输入。本文提出一种新方法——视频-3D几何大语言模型(VG LLM),仅通过视频序列提取3D几何先验信息,将其与视觉特征融合后输入多模态大语言模型(MLLM),实现对3D场景的直接理解与空间推理。大量实验表明,该方法在多种3D场景理解与空间推理任务中均有显著提升,全部基于视频数据训练。令人印象深刻的是,仅使用40亿参数的模型,在不依赖显式3D数据的情况下,性能媲美甚至超过当前最优方法,尤其在VSI-Bench评测中优于Gemini-1.5-Pro。
原文摘要 · Abstract (English)
Previous research has investigated the application of Multimodal Large Language Models (MLLMs) in understanding 3D scenes by interpreting them as videos. These approaches generally depend on comprehensive 3D data inputs, such as point clouds or reconstructed Bird's-Eye View (BEV) maps. In our research, we advance this field by enhancing the capability of MLLMs to understand and reason in 3D spaces directly from video data, without the need for additional 3D input. We propose a novel and efficient method called the Video-3D Geometry Large Language Model (VG LLM). Our approach utilizes a 3D visual geometry encoder to extract 3D prior information from video sequences. This information is then integrated with visual tokens and input into the MLLM. Extensive experiments have shown that our method has achieved substantial improvements in various tasks related to 3D scene understanding and spatial reasoning, all directly learned from video sources. Impressively, our 4B model, which does not rely on explicit 3D data inputs, achieves competitive results compared to existing state-of-the-art methods, and even surpasses the Gemini-1.5-Pro in the VSI-Bench evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。