arXiv:2605.30307cs.CV2026-05

让视觉语言模型实时引用图像区域并推断3D空间位置,提升空间理解能力。

Grounded 3D-Aware Spatial Vision-Language Modeling

论文配图:Grounded 3D-Aware Spatial Vision-Language Modeling
图 1 · 摘自论文原文
  • 引入隐式定位机制,在生成时动态插入视觉区域标记。
  • 通过单目图像预测3D边界框,精度优于基线模型12.7%。
  • 适合需要精准空间推理的多模态任务,如机器人导航与智能驾驶。

我们提出GR3D,一种具备三种互补定位能力的时空视觉语言模型:显式2D定位、隐式2D定位和单目3D定位。GR3D引入隐式定位机制,在生成过程中识别实体提及并插入对应区域标记,使模型可在生成空间推理链时即时引用视觉证据。同时,采用区域提示的单目3D定位设计,从定位区域查询中预测相机视角下的3D边界框,结合内在感知归一化和密集几何监督。这些定位能力共同实现从2D感知到3D推理的分解式空间理解。GR3D在多种有无定位的空间基准测试中均取得稳定提升,证明定位作为归纳偏置对强化视觉语言模型空间理解的有效性。这些能力显著提升了模型在定位任务之外的通用空间认知能力。

原文摘要 · Abstract (English)

We present GR3D, a spatial vision language model equipped with three complementary grounding capabilities--explicit 2D grounding, implicit 2D grounding, and monocular 3D grounding--within a single framework. GR3D introduces an implicit grounding mechanism that identifies entity mentions during generation and inserts the corresponding region tokens into the text stream, allowing the model to reference visual evidence on the fly when producing spatial chain-of-thought responses. In parallel, a region-prompted monocular 3D grounding design predicts 3D bounding boxes in the camera view from grounded region queries, supported by intrinsic-aware normalization and dense geometric supervision. Together, these grounding capabilities enable GR3D to decompose complex spatial understanding problems into grounded 2D perception followed by 3D inference. GR3D achieves consistent improvements across grounded and non-grounded spatial benchmarks, demonstrating grounding as an effective inductive bias for strengthening spatial understanding in VLMs. These grounding capabilities collectively enhance general spatial understanding beyond the grounding task itself.

空间理解视觉语言模型3D推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。