arXiv:2603.15386cs.CVcs.AI2026-03被引 3

用三维几何图解构场景,让大模型更准地推理空间关系。

RieMind: Geometry-Grounded Spatial Agent for Scene Understanding

  • 把场景转成带几何信息的3D图谱,让大模型通过结构化工具推理。
  • 在理想感知条件下,空间推理准确率比之前方法高16%。
  • 适合研究视觉推理、空间认知或想突破纯视觉模型的学者。

视觉语言模型(VLMs)已成为理解室内场景的主流范式,但在度量与空间推理方面仍存在不足。现有方法依赖端到端视频理解或大规模空间问答微调,将感知与推理耦合。本文探讨解耦感知与推理是否能提升空间推理能力。提出一种针对静态3D室内场景的代理框架,将大模型嵌入显式的三维场景图(3DSG)中。每个场景由专用感知模块构建为持久的3DSG,基于真实标注数据。代理仅通过结构化几何工具与场景交互,暴露物体尺寸、距离、位姿及空间关系等基础属性。在VSI-Bench的静态测试集上,我们获得理想感知条件下的性能上限,结果显著优于此前工作,最高提升达16%,且无需任务特定微调。相比基线VLMs,本方案平均提升33%至50%。结果表明,显式几何接地可大幅提升空间推理表现,提示结构化表示是纯端到端视觉推理的有力替代方案。

原文摘要 · Abstract (English)

Visual Language Models (VLMs) have increasingly become the main paradigm for understanding indoor scenes, but they still struggle with metric and spatial reasoning. Current approaches rely on end-to-end video understanding or large-scale spatial question answering fine-tuning, inherently coupling perception and reasoning. In this paper, we investigate whether decoupling perception and reasoning leads to improved spatial reasoning. We propose an agentic framework for static 3D indoor scene reasoning that grounds an LLM in an explicit 3D scene graph (3DSG). Rather than ingesting videos directly, each scene is represented as a persistent 3DSG constructed by a dedicated perception module. To isolate reasoning performance, we instantiate the 3DSG from ground-truth annotations. The agent interacts with the scene exclusively through structured geometric tools that expose fundamental properties such as object dimensions, distances, poses, and spatial relationships. The results we obtain on the static split of VSI-Bench provide an upper bound under ideal perceptual conditions on the spatial reasoning performance, and we find that it is significantly higher than previous works, by up to 16\%, without task specific fine-tuning. Compared to base VLMs, our agentic variant achieves significantly better performance, with average improvements between 33\% to 50\%. These findings indicate that explicit geometric grounding substantially improves spatial reasoning performance, and suggest that structured representations offer a compelling alternative to purely end-to-end visual reasoning.

空间推理三维场景图智能体几何建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。