自动为3D物体检测生成类别名,无需人工指定
Auto-Vocabulary 3D Object Detection
- 用2D视觉语言模型自动生成语义候选类别
- 在ScanNetV2上实现3.48%的mAP提升和24.5%的语义质量改善
- 适合开放词汇3D检测与自动化标注场景
开放词汇3D目标检测方法可定位训练时未见类别的3D框。尽管名称如此,现有方法仍需用户在训练和推理阶段指定类别。我们提出自动词汇3D目标检测(AV3DOD),在无需任何用户输入的情况下自动为检测到的物体生成类别名。为此,我们引入语义得分(SS)评估生成类别名的质量。进一步提出新框架AV3DOD,利用2D视觉语言模型通过图像描述、伪3D框生成和特征空间语义扩展生成丰富语义候选。AV3DOD在ScanNetV2和SUNRGB-D数据集上均达到最新性能,在定位(mAP)和语义质量(SS)方面表现优异。显著优于当前最优方法CoDA:在ScanNetV2上整体mAP提升3.48%,语义得分相对提高24.5%。
原文摘要 · Abstract (English)
Open-vocabulary 3D object detection methods are able to localize 3D boxes of classes unseen during training. Despite the name, existing methods rely on user-specified classes both at training and inference. We propose to study Auto-Vocabulary 3D Object Detection (AV3DOD), where the classes are automatically generated for the detected objects without any user input. To this end, we introduce Semantic Score (SS) to evaluate the quality of the generated class names. We then develop a novel framework, AV3DOD, which leverages 2D vision-language models (VLMs) to generate rich semantic candidates through image captioning, pseudo 3D box generation, and feature-space semantics expansion. AV3DOD achieves the state-of-the-art (SOTA) performance on both localization (mAP) and semantic quality (SS) on the ScanNetV2 and SUNRGB-D datasets. Notably, it surpasses the SOTA, CoDA, by 3.48 overall mAP and attains a 24.5% relative improvement in SS on ScanNetV2.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。