用大模型指导3D特征分组,实现遮挡下的开放词汇视觉定位。
ReasonGrounder: LVLM-Guided Hierarchical Feature Splatting for Open-Vocabulary 3D Visual Grounding and Reasoning
- 基于大视觉语言模型解析指令,按物理尺度自适应分组3D高斯特征
- 在10万+场景上提升遮挡物体定位准确率,支持隐含语义理解
- 适合机器人导航、自动驾驶等需推理的开放世界任务
开放词汇3D视觉定位与推理旨在根据隐含语言描述定位场景中的物体,即使物体被遮挡。这对视觉语言导航和自主机器人至关重要。然而,现有方法依赖大量3D标注与掩码提议进行微调,难以处理多样语义和常识推理。本文提出ReasonGrounder,一种基于大视觉语言模型(LVLM)的框架,利用分层3D特征高斯场实现基于物理尺度的自适应分组,支持开放词汇3D定位与推理。该方法通过大视觉语言模型解析隐含指令,借助3D高斯溅射定位遮挡物体。结合SAM提供的2D分割掩码与多视角CLIP嵌入,根据物体尺度选择高斯组,实现对显性与隐性语言的理解,在新视图与遮挡情况下仍能精准定位。我们还构建了ReasoningGD数据集,包含超过10,000个场景和200万条标注,用于评估开放词汇3D定位与非完整感知能力。实验表明,ReasonGrounder在真实场景中显著提升了3D定位准确率。
原文摘要 · Abstract (English)
Open-vocabulary 3D visual grounding and reasoning aim to localize objects in a scene based on implicit language descriptions, even when they are occluded. This ability is crucial for tasks such as vision-language navigation and autonomous robotics. However, current methods struggle because they rely heavily on fine-tuning with 3D annotations and mask proposals, which limits their ability to handle diverse semantics and common knowledge required for effective reasoning. In this work, we propose ReasonGrounder, an LVLM-guided framework that uses hierarchical 3D feature Gaussian fields for adaptive grouping based on physical scale, enabling open-vocabulary 3D grounding and reasoning. ReasonGrounder interprets implicit instructions using large vision-language models (LVLM) and localizes occluded objects through 3D Gaussian splatting. By incorporating 2D segmentation masks from the SAM and multi-view CLIP embeddings, ReasonGrounder selects Gaussian groups based on object scale, enabling accurate localization through both explicit and implicit language understanding, even in novel, occluded views. We also contribute ReasoningGD, a new dataset containing over 10K scenes and 2 million annotations for evaluating open-vocabulary 3D grounding and amodal perception under occlusion. Experiments show that ReasonGrounder significantly improves 3D grounding accuracy in real-world scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。