通过分层对齐实现3D目标检测的开放词汇识别,提升新物体识别能力。
Hierarchical Cross-Modal Alignment for Open-Vocabulary 3D Object Detection
- 设计分层数据融合方法,从粗到细构建3D图像文本数据
- 在多个基准上超越现有最优方法,无3D标注也能取得良好效果
- 适合需要识别未知物体的自动驾驶与机器人场景
开放词汇3D目标检测旨在定位并分类超出封闭词表的新物体。视觉语言模型(VLM)在理解开放词汇方面展现出强大能力。现有基于VLM的3D目标检测方法通常丢失3D感知所需的丰富场景上下文。为此,本文提出一种分层框架HCMA,同时学习局部物体与全局场景信息。首先设计分层数据集成(HDI)方法,生成从粗到细的3D-图像-文本数据,并输入VLM提取以物体为中心的知识。为促进特征层级间的关联,提出交互式跨模态对齐(ICMA)策略,建立层内与层间有效连接。进一步设计物体聚焦上下文调节(OFCA)模块,通过强化物体相关特征来优化多层级特征。大量实验表明,所提方法在现有开放词汇3D检测基准上优于当前最优方法,且在无任何3D标注情况下仍取得有竞争力的结果。
原文摘要 · Abstract (English)
Open-vocabulary 3D object detection (OV-3DOD) aims at localizing and classifying novel objects beyond closed sets. The recent success of vision-language models (VLMs) has demonstrated their remarkable capabilities to understand open vocabularies. Existing works that leverage VLMs for 3D object detection (3DOD) generally resort to representations that lose the rich scene context required for 3D perception. To address this problem, we propose in this paper a hierarchical framework, named HCMA, to simultaneously learn local object and global scene information for OV-3DOD. Specifically, we first design a Hierarchical Data Integration (HDI) approach to obtain coarse-to-fine 3D-image-text data, which is fed into a VLM to extract object-centric knowledge. To facilitate the association of feature hierarchies, we then propose an Interactive Cross-Modal Alignment (ICMA) strategy to establish effective intra-level and inter-level feature connections. To better align features across different levels, we further propose an Object-Focusing Context Adjustment (OFCA) module to refine multi-level features by emphasizing object-related features. Extensive experiments demonstrate that the proposed method outperforms SOTA methods on the existing OV-3DOD benchmarks. It also achieves promising OV-3DOD results even without any 3D annotations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。