arXiv:2603.18002cs.CVcs.AI2026-03被引 1

让视觉语言模型从2D图像理解3D空间,实现精准定位与场景推理。

Loc3R-VLM: Language-based Localization and 3D Reasoning with Vision-Language Models

  • 通过重建全局布局和建模视角位置,实现3D空间感知。
  • 在多个3D问答任务上超越现有方法,提升定位准确率。
  • 适合需要三维理解的机器人导航与智能交互研究者。

多模态大语言模型在视觉与语言关联方面进展显著,但在空间理解与视角感知方面仍存不足。现有工作通常通过引入几何线索增强输入表示,而非显式训练模型进行3D推理。本文提出Loc3R-VLM框架,将单目视频输入下的2D视觉语言模型升级为具备先进3D理解能力的系统。受人类空间认知启发,该框架采用双目标联合优化:全局布局重建以构建场景结构的整体表征,显式情境建模以锚定自身视角。这两个目标提供直接的空间监督,使感知与语言均扎根于3D语境。为确保几何一致性与度量尺度对齐,利用预训练3D基础模型提取轻量级相机位姿先验。实验表明,Loc3R-VLM在语言驱动定位任务中达到当前最优性能,并在情境化与通用3D问答基准上优于现有2D及视频基方法,验证了其空间监督框架对强3D理解的有效性。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have made impressive progress in connecting vision and language, but they still struggle with spatial understanding and viewpoint-aware reasoning. Recent efforts aim to augment the input representations with geometric cues rather than explicitly teaching models to reason in 3D space. We introduce Loc3R-VLM, a framework that equips 2D Vision-Language Models with advanced 3D understanding capabilities from monocular video input. Inspired by human spatial cognition, Loc3R-VLM relies on two joint objectives: global layout reconstruction to build a holistic representation of the scene structure, and explicit situation modeling to anchor egocentric perspective. These objectives provide direct spatial supervision that grounds both perception and language in a 3D context. To ensure geometric consistency and metric-scale alignment, we leverage lightweight camera pose priors extracted from a pre-trained 3D foundation model. Loc3R-VLM achieves state-of-the-art performance in language-based localization and outperforms existing 2D- and video-based approaches on situated and general 3D question-answering benchmarks, demonstrating that our spatial supervision framework enables strong 3D understanding. Project page: https://kevinqu7.github.io/loc3r-vlm

3D理解视觉语言模型空间推理定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。