arXiv:2511.21688cs.CVcs.AI2025-11被引 31

让视觉语言模型学会三维空间重建与推理,提升空间理解能力。

G$^2$VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning

  • 用2D图像学习3D几何特征,统一建模空间重建与理解。
  • 在空间推理任务上表现优于或媲美顶尖3D重建模型。
  • 适合做3D场景编辑、智能机器人导航等空间智能应用。

视觉语言模型在空间理解与推理任务中仍表现不佳,我们归因于缺乏从2D图像重建3D空间的视觉几何学习能力。本文提出G²VLM,一种基于几何感知的视觉语言模型,统一实现三维空间重建与空间理解。该模型原生利用学习到的3D视觉几何特征,通过上下文学习和交错推理,直接预测3D属性并增强空间推理能力。其统一架构可扩展性强:训练依赖大量多视角图像与视频数据,同时融合通常需难获取标注的3D先验知识。实验表明,G²VLM在两项任务上均表现优异,其三维重建性能可媲美当前最优前馈模型,在多种空间理解与推理任务中达到更优或竞争力结果。通过结合语义强的VLM与底层3D视觉任务,我们期望G²VLM能成为社区基准,推动3D场景编辑等未来应用发展。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) still lack robustness in spatial intelligence, demonstrating poor performance on spatial understanding and reasoning tasks. We attribute this gap to the absence of a visual geometry learning process capable of reconstructing 3D space from 2D images. We present G$^2$VLM, a geometry grounded vision-language model that bridges two fundamental aspects of spatial intelligence: spatial 3D reconstruction and spatial understanding. G$^2$VLM natively leverages learned 3D visual geometry features to directly predict 3D attributes and enhance spatial reasoning tasks via in-context learning and interleaved reasoning. Our unified design is highly scalable for spatial understanding: it trains on abundant multi-view image and video data, while simultaneously leveraging the benefits of 3D visual priors that are typically only derived from hard-to-collect annotations. Experimental results demonstrate G$^2$VLM is proficient in both tasks, achieving comparable results to state-of-the-art feed-forward 3D reconstruction models and achieving better or competitive results across spatial understanding and reasoning tasks. By unifying a semantically strong VLM with low-level 3D vision tasks, we hope G$^2$VLM can serve as a strong baseline for the community and unlock more future applications, such as 3D scene editing.

视觉语言模型三维重建空间推理几何感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。