让视觉语言模型学会三维空间重建与推理,提升空间理解能力。
G$^2$VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning
- 用2D图像学习3D几何特征,统一建模空间重建与理解。
- 在空间推理任务上表现优于或媲美顶尖3D重建模型。
- 适合做3D场景编辑、智能机器人导航等空间智能应用。
视觉语言模型在空间理解与推理任务中仍表现不佳,我们归因于缺乏从2D图像重建3D空间的视觉几何学习能力。本文提出G²VLM,一种基于几何感知的视觉语言模型,统一实现三维空间重建与空间理解。该模型原生利用学习到的3D视觉几何特征,通过上下文学习和交错推理,直接预测3D属性并增强空间推理能力。其统一架构可扩展性强:训练依赖大量多视角图像与视频数据,同时融合通常需难获取标注的3D先验知识。实验表明,G²VLM在两项任务上均表现优异,其三维重建性能可媲美当前最优前馈模型,在多种空间理解与推理任务中达到更优或竞争力结果。通过结合语义强的VLM与底层3D视觉任务,我们期望G²VLM能成为社区基准,推动3D场景编辑等未来应用发展。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) still lack robustness in spatial intelligence, demonstrating poor performance on spatial understanding and reasoning tasks. We attribute this gap to the absence of a visual geometry learning process capable of reconstructing 3D space from 2D images. We present G$^2$VLM, a geometry grounded vision-language model that bridges two fundamental aspects of spatial intelligence: spatial 3D reconstruction and spatial understanding. G$^2$VLM natively leverages learned 3D visual geometry features to directly predict 3D attributes and enhance spatial reasoning tasks via in-context learning and interleaved reasoning. Our unified design is highly scalable for spatial understanding: it trains on abundant multi-view image and video data, while simultaneously leveraging the benefits of 3D visual priors that are typically only derived from hard-to-collect annotations. Experimental results demonstrate G$^2$VLM is proficient in both tasks, achieving comparable results to state-of-the-art feed-forward 3D reconstruction models and achieving better or competitive results across spatial understanding and reasoning tasks. By unifying a semantically strong VLM with low-level 3D vision tasks, we hope G$^2$VLM can serve as a strong baseline for the community and unlock more future applications, such as 3D scene editing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。