arXiv:2503.22436cs.CV2025-03被引 14

提出首个自动驾驶多视角3D视觉定位基准,提升语言指令理解与目标定位精度。

NuGrounding: A Multi-View 3D Visual Grounding Framework in Autonomous Driving

  • 构建分层语言指令生成框架,覆盖人类真实指令模式。
  • 融合多模态大模型与检测模型,实现0.59精度与0.64召回率。
  • 适合研究自动驾驶语义感知与跨模态对齐的学者参考。

多视角3D视觉定位对自动驾驶车辆理解自然语言并定位复杂环境中的目标物体至关重要。然而,现有数据集和方法存在语言指令粗粒度、三维几何推理与语言理解融合不足的问题。为此,我们提出NuGrounding,首个面向自动驾驶的多视角3D视觉定位大规模基准。通过层级化定位(HoG)方法生成多层次指令,全面覆盖人类指令模式。针对该挑战性数据集,我们提出一种新范式,无缝结合多模态大模型(MLLMs)的语言理解能力与专用检测模型的精确定位能力。方法引入两个解耦任务标记和一个上下文查询,聚合三维几何信息与语义指令,再经融合解码器优化空间-语义特征融合,实现精准定位。大量实验表明,本方法显著优于现有代表性3D场景理解基线,在精度上达0.59,召回率达0.64,分别提升50.8%和54.7%。

原文摘要 · Abstract (English)

Multi-view 3D visual grounding is critical for autonomous driving vehicles to interpret natural languages and localize target objects in complex environments. However, existing datasets and methods suffer from coarse-grained language instructions, and inadequate integration of 3D geometric reasoning with linguistic comprehension. To this end, we introduce NuGrounding, the first large-scale benchmark for multi-view 3D visual grounding in autonomous driving. We present a Hierarchy of Grounding (HoG) method to construct NuGrounding to generate hierarchical multi-level instructions, ensuring comprehensive coverage of human instruction patterns. To tackle this challenging dataset, we propose a novel paradigm that seamlessly combines instruction comprehension abilities of multi-modal LLMs (MLLMs) with precise localization abilities of specialist detection models. Our approach introduces two decoupled task tokens and a context query to aggregate 3D geometric information and semantic instructions, followed by a fusion decoder to refine spatial-semantic feature fusion for precise localization. Extensive experiments demonstrate that our method significantly outperforms the baselines adapted from representative 3D scene understanding methods by a significant margin and achieves 0.59 in precision and 0.64 in recall, with improvements of 50.8% and 54.7%.

3D视觉定位自动驾驶多模态语言理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。