用视觉语言模型提升3D定位的语义理解与视角推理能力
GuideGround: VLM-guided Semantic Understanding and Viewpoint-aware Reasoning for 3D Visual Grounding

- 用VLM生成开放词汇语义描述替代封闭分类,增强语义理解
- 保留各视角独立定位假设,通过VLM跨视角验证提升准确性
- 在ReferIt3D上超越当前最佳方法,适合需要精准3D定位的研究
3D视觉定位旨在从自然语言查询中定位3D场景中的目标物体,需兼具细粒度语义理解与视角依赖的空间推理。现有方法通常将语义理解建模为辅助的封闭集物体分类任务,并依赖多视角特征聚合进行视角推理,限制了语义泛化能力并削弱了视角特定证据。我们观察到,视觉语言模型(VLM)天然具备开放词汇语义理解与全局场景感知能力。基于此,我们提出GuideGround,一种借助VLM增强而非取代任务专用定位模型的框架,通过VLM实现语义增强和视角特定假设验证。具体而言,我们以VLM生成的物体语义描述替代辅助的封闭集分类任务,以增强语义理解;同时,不直接聚合多视角表征,而是保留各视角的定位假设,并显式利用VLM在候选视角间验证这些假设。在ReferIt3D基准上的大量实验表明,GuideGround持续优于先前最先进方法。全面的消融研究进一步证实了所提语义理解与视角推理策略的有效性。
原文摘要 · Abstract (English)
3D visual grounding aims to localize the target object in a 3D scene from a natural language query, requiring both fine-grained semantic understanding and viewpoint-dependent spatial reasoning. Existing methods typically formulate semantic understanding as an auxiliary closed-set object classification task and rely on multi-view feature aggregation for viewpoint reasoning, limiting semantic generalization and weakening viewpoint-specific evidence. We observe that vision-language models naturally provide complementary capabilities through open-vocabulary semantic understanding and global scene perception. Based on this insight, we propose GuideGround, a VLM-guided framework that complements rather than replaces task-specific grounding models by leveraging VLMs for semantic enhancement and viewpoint-specific hypothesis verification. Specifically, we replace auxiliary closed-set object classification with VLM-generated object semantic descriptions to enhance semantic understanding. Meanwhile, instead of directly aggregating multi-view representations, we preserve viewpoint-specific grounding hypotheses through per-view grounding and explicitly verify them using VLMs across candidate viewpoints. Extensive experiments on the ReferIt3D benchmark demonstrate that GuideGround consistently outperforms previous state-of-the-art methods. Comprehensive ablation studies further confirm the effectiveness of both the proposed semantic understanding and viewpoint reasoning strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。