arXiv:2412.17800cs.CV2024-12被引 2

用多模态原型提升海量类别目标检测,简单有效。

Comprehensive Multi-Modal Prototypes are Simple and Effective Classifiers for Vast-Vocabulary Object Detection

  • 用视觉语言模型生成全面的多模态原型作为分类器初始
  • 在V3Det上使Faster R-CNN等模型性能提升3.3~6.2个点
  • 适合需要大规模开放词汇检测的应用场景

实现对海量开放世界类别的目标检测是长期追求的目标。借助视觉语言模型的泛化能力,当前开放词汇检测器虽仅在有限类别上训练,仍能识别更广范围的词汇。然而,当训练时类别词表规模扩展至真实世界水平,以往基于粗粒度类别名对齐的分类器显著降低检测性能。本文提出Prova,一种用于海量词汇目标检测的多模态原型分类器。Prova通过提取全面的多模态原型作为对齐分类器的初始化,解决海量词汇识别失败问题。在V3Det数据集上,该方法仅增加投影层,即可显著提升单阶段、两阶段及DETR-based检测器性能。在监督设置下,Faster R-CNN、FCOS和DINO分别提升3.3、6.2和2.9个AP;在开放词汇设置下,达到32.8基类AP和11.0新类AP,分别超越此前方法2.6和4.3个点,刷新性能纪录。

原文摘要 · Abstract (English)

Enabling models to recognize vast open-world categories has been a longstanding pursuit in object detection. By leveraging the generalization capabilities of vision-language models, current open-world detectors can recognize a broader range of vocabularies, despite being trained on limited categories. However, when the scale of the category vocabularies during training expands to a real-world level, previous classifiers aligned with coarse class names significantly reduce the recognition performance of these detectors. In this paper, we introduce Prova, a multi-modal prototype classifier for vast-vocabulary object detection. Prova extracts comprehensive multi-modal prototypes as initialization of alignment classifiers to tackle the vast-vocabulary object recognition failure problem. On V3Det, this simple method greatly enhances the performance among one-stage, two-stage, and DETR-based detectors with only additional projection layers in both supervised and open-vocabulary settings. In particular, Prova improves Faster R-CNN, FCOS, and DINO by 3.3, 6.2, and 2.9 AP respectively in the supervised setting of V3Det. For the open-vocabulary setting, Prova achieves a new state-of-the-art performance with 32.8 base AP and 11.0 novel AP, which is of 2.6 and 4.3 gain over the previous methods.

目标检测多模态开放词汇原型学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。