arXiv:2507.06719cs.CVcs.RO2025-07被引 8

用大模型增强空间推理,让AI更准理解语言中的位置关系。

A Neural Representation Framework with LLM-Driven Spatial Reasoning for Open-Vocabulary 3D Visual Grounding

  • 用大模型解析语言中的空间关系,明确目标、参照物和相对位置。
  • 结合透明度和颜色构建分层特征场,提升3D场景理解精度。
  • 可嵌入多种神经表示框架,适合机器人导航等实际应用。

开放词汇3D视觉定位旨在根据自由形式的语言查询定位目标物体,对自主导航、机器人和增强现实等具身智能应用至关重要。通过神经表示学习3D语言场,可从有限视角准确理解3D场景并定位复杂环境中的目标。然而现有方法在处理语言查询中的空间关系(如“椅子上的书”)时表现不佳,主要源于对语言与3D场景中空间关系的推理不足。本文提出SpatialReasoner——一种基于神经表示、由大语言模型驱动空间推理的新框架,构建融合视觉属性的分层特征场以实现开放词汇3D视觉定位。为实现语言中的空间推理,该框架微调大模型以捕捉空间关系,并显式推断目标、锚点及空间关系指令;为实现3D场景中的空间推理,引入不透明度与颜色信息,利用蒸馏的CLIP特征和来自Segment Anything Model(SAM)提取的掩码构建分层特征场。该场通过分层查询方式,依据语言查询中的空间关系定位目标3D实例。大量实验表明,本框架可无缝集成于不同神经表示,显著优于基线模型,并赋予其空间推理能力。

原文摘要 · Abstract (English)

Open-vocabulary 3D visual grounding aims to localize target objects based on free-form language queries, which is crucial for embodied AI applications such as autonomous navigation, robotics, and augmented reality. Learning 3D language fields through neural representations enables accurate understanding of 3D scenes from limited viewpoints and facilitates the localization of target objects in complex environments. However, existing language field methods struggle to accurately localize instances using spatial relations in language queries, such as ``the book on the chair.'' This limitation mainly arises from inadequate reasoning about spatial relations in both language queries and 3D scenes. In this work, we propose SpatialReasoner, a novel neural representation-based framework with large language model (LLM)-driven spatial reasoning that constructs a visual properties-enhanced hierarchical feature field for open-vocabulary 3D visual grounding. To enable spatial reasoning in language queries, SpatialReasoner fine-tunes an LLM to capture spatial relations and explicitly infer instructions for the target, anchor, and spatial relation. To enable spatial reasoning in 3D scenes, SpatialReasoner incorporates visual properties (opacity and color) to construct a hierarchical feature field. This field represents language and instance features using distilled CLIP features and masks extracted via the Segment Anything Model (SAM). The field is then queried using the inferred instructions in a hierarchical manner to localize the target 3D instance based on the spatial relation in the language query. Extensive experiments show that our framework can be seamlessly integrated into different neural representations, outperforming baseline models in 3D visual grounding while empowering their spatial reasoning capability.

3D视觉空间推理大模型开放词汇

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。