arXiv:2505.04965cs.CV2025-05ICLR被引 14

提升第一人称3D视觉定位的语义精度,让智能体更懂自然语言描述。

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding

  • 用分层场景增强器保留点云与图像的密集视觉语义
  • 训练时用大模型生成丰富语言描述,提升文本上下文
  • 在多个数据集上超越现有方法,获CVPR 2024创新奖

让智能体通过自然语言理解并交互于3D环境是推动机器人和人机交互的关键。核心任务是第一人称3D视觉定位,即根据语言描述在真实3D空间中定位目标物体。当前面临两大挑战:(1) 点云与第一人称多视角图像稀疏融合导致细粒度视觉语义丢失;(2) 语言描述任意性导致文本语义上下文有限。我们提出DenseGrounding,通过增强视觉与文本语义来解决上述问题。视觉方面,引入分层场景语义增强器,捕获细粒度全局场景特征并促进跨模态对齐;文本方面,提出语言语义增强器,利用大语言模型在训练时提供丰富上下文与多样化语言描述。大量实验表明,DenseGrounding在整体准确率上显著优于现有方法,全数据集训练下提升5.81%,小样本子集上提升7.56%,进一步推进了该领域的最先进水平。本方法还在CVPR 2024自动驾驶大奖赛多视图3D视觉定位赛道中获得第一名,并荣获创新奖,验证其有效性与鲁棒性。

原文摘要 · Abstract (English)

Enabling intelligent agents to comprehend and interact with 3D environments through natural language is crucial for advancing robotics and human-computer interaction. A fundamental task in this field is ego-centric 3D visual grounding, where agents locate target objects in real-world 3D spaces based on verbal descriptions. However, this task faces two significant challenges: (1) loss of fine-grained visual semantics due to sparse fusion of point clouds with ego-centric multi-view images, (2) limited textual semantic context due to arbitrary language descriptions. We propose DenseGrounding, a novel approach designed to address these issues by enhancing both visual and textual semantics. For visual features, we introduce the Hierarchical Scene Semantic Enhancer, which retains dense semantics by capturing fine-grained global scene features and facilitating cross-modal alignment. For text descriptions, we propose a Language Semantic Enhancer that leverages large language models to provide rich context and diverse language descriptions with additional context during model training. Extensive experiments show that DenseGrounding significantly outperforms existing methods in overall accuracy, with improvements of 5.81% and 7.56% when trained on the comprehensive full dataset and smaller mini subset, respectively, further advancing the SOTA in egocentric 3D visual grounding. Our method also achieves 1st place and receives the Innovation Award in the CVPR 2024 Autonomous Grand Challenge Multi-view 3D Visual Grounding Track, validating its effectiveness and robustness.

3D视觉定位多模态大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。