arXiv:2601.13132cs.CV2026-01中稿 · ECCV被引 1

通过生成新视角增强视觉语言模型对三维场景的理解与定位能力

SplatReasoner: Enhancing Embodied Reasoning and Grounding by Novel View Synthesis

  • 用3D高斯点阵合成查询相关的新增视角
  • 在三维推理和实体定位任务上显著优于固定视角基线
  • 适合需要精准空间理解的机器人导航与交互场景

视觉语言模型在图像和视频推理上表现优异,但在具身场景理解中常受限于存储在情景式RGB-D记忆中的固定视角。这些视角可能因遮挡、物体截断、视野受限或视角不佳而遗漏查询相关证据。我们提出SplatReasoner框架,通过3D高斯点阵(3DGS)将新视角合成引入视觉语言模型的推理过程。针对用户对三维场景的查询,SplatReasoner检索相关观测,并生成条件化的新增视角,以揭示回答查询所需的视觉证据并实现对指称实体的三维定位。实验表明,查询条件下的新视角合成在具身推理和三维定位性能上均优于固定视角记忆和嵌入语言的3DGS基线。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) have demonstrated strong reasoning capabilities over images and videos, yet their application to embodied scene understanding often constrained by the fixed viewpoints stored in episodic RGB-D memories. These observations may fail to capture query-relevant evidence due to occlusions, object truncation, restricted fields of view, or suboptimal view composition. We present SplatReasoner, a framework that introduces novel view synthesis into the VLM reasoning process by leveraging 3D Gaussian Splatting (3DGS). Given a user query about a 3D scene, SplatReasoner retrieves relevant observations and synthesizes query-conditioned viewpoints that reveal the visual evidence needed to answer the query and ground the referred entities in 3D. Experiments show that query-conditioned novel view synthesis improves both embodied reasoning and 3D grounding over fixed-view memory and language-embedded 3DGS baselines.

具身智能三维理解新视角合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。