提出3D空间推理新框架,让模型像人一样理解物体位置与关系。
Reasoning in Space via Grounding in the World
- 设计双路池化机制,统一融合语义、几何与位置信息。
- 无需外部模块实现端到端3D定位,性能达顶尖水平。
- 构建带推理链的标注数据集,推动视觉与空间推理融合。
本文认为3D视觉定位是空间推理的核心,提出Grounded-Spatial Reasoner(GS-Reasoner)以构建连接两者的统一3D表示。现有3D大模型缺乏能同时捕捉语义与几何信息的统一表征,导致定位能力差或过度依赖外部模块,阻碍了定位与空间推理的无缝整合。为此,我们提出一种简洁有效的双路池化机制,将几何特征紧密对齐语义与位置线索,构建基于图像块的统一3D表示,不增加输入令牌数。基于此整体表征,GS-Reasoner首次实现完全无需外部模块的自回归定位,性能媲美最先进模型,建立统一自洽的3D空间推理框架。为进一步弥合定位与空间推理的差距,我们引入了包含3D边界框标注和逐步推理路径的Grounded Chain-of-Thought(GCoT)数据集。大量实验表明,GS-Reasoner在3D视觉定位上表现优异,显著提升空间推理能力,达到当前最优水平。
原文摘要 · Abstract (English)
In this paper, we claim that 3D visual grounding is the cornerstone of spatial reasoning and introduce the Grounded-Spatial Reasoner (GS-Reasoner) to explore the effective spatial representations that bridge the gap between them. Existing 3D LLMs suffer from the absence of a unified 3D representation capable of jointly capturing semantic and geometric information. This deficiency is manifested either in poor performance on grounding or in an excessive reliance on external modules, ultimately hindering the seamless integration of grounding and spatial reasoning. To address this, we propose a simple yet effective dual-path pooling mechanism that tightly aligns geometric features with both semantic and positional cues, constructing a unified image patch-based 3D representation that encapsulates all essential information without increasing the number of input tokens. Leveraging this holistic representation, GS-Reasoner is the first 3D LLM that achieves autoregressive grounding entirely without external modules while delivering performance comparable to state-of-the-art models, establishing a unified and self-contained framework for 3D spatial reasoning. To further bridge grounding and spatial reasoning, we introduce the Grounded Chain-of-Thought (GCoT) dataset. This dataset is meticulously curated to include both 3D bounding box annotations for objects referenced in reasoning questions and step-by-step reasoning paths that integrate grounding as a core component of the problem-solving process. Extensive experiments demonstrate that GS-Reasoner achieves impressive results on 3D visual grounding, which in turn significantly enhances its spatial reasoning capabilities, leading to state-of-the-art performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。