arXiv:2602.00531cs.CV2026-02被引 1

通过多层级视觉语言对齐,提升开放词汇目标检测性能。

Enhancing Open-Vocabulary Object Detection through Multi-Level Fine-Grained Visual-Language Alignment

  • 构建特征金字塔增强细粒度视觉语言对齐。
  • 在COCO2017上新类别达58.7 AP,LVIS上达24.8 AP。
  • 适合需要识别未见类别的实际场景应用。

传统目标检测系统受限于预定义类别,难以适应动态环境。开放词汇目标检测(OVD)可识别训练集外的新类别。现有方法或难以将CLIP单尺度图像主干适配至检测框架,或缺乏稳健的视觉-语言对齐。本文提出视觉语言检测框架VLDet,通过重构特征金字塔实现细粒度对齐。引入VL-PUB模块,有效利用CLIP的视觉-语言知识,并通过特征金字塔适配检测任务。此外,设计SigRPN块,采用基于Sigmoid的锚框-文本对比损失,提升新类别检测能力。大量实验表明,该方法在COCO2017上新类别达58.7 AP,LVIS上达24.8 AP,分别优于当前最优方法27.6%和6.9%。同时在封闭集检测中也展现优异零样本性能。

原文摘要 · Abstract (English)

Traditional object detection systems are typically constrained to predefined categories, limiting their applicability in dynamic environments. In contrast, open-vocabulary object detection (OVD) enables the identification of objects from novel classes not present in the training set. Recent advances in visual-language modeling have led to significant progress of OVD. However, prior works face challenges in either adapting the single-scale image backbone from CLIP to the detection framework or ensuring robust visual-language alignment. We propose Visual-Language Detection (VLDet), a novel framework that revamps feature pyramid for fine-grained visual-language alignment, leading to improved OVD performance. With the VL-PUB module, VLDet effectively exploits the visual-language knowledge from CLIP and adapts the backbone for object detection through feature pyramid. In addition, we introduce the SigRPN block, which incorporates a sigmoid-based anchor-text contrastive alignment loss to improve detection of novel categories. Through extensive experiments, our approach achieves 58.7 AP for novel classes on COCO2017 and 24.8 AP on LVIS, surpassing all state-of-the-art methods and achieving significant improvements of 27.6% and 6.9%, respectively. Furthermore, VLDet also demonstrates superior zero-shot performance on closed-set object detection.

目标检测开放词汇视觉语言对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。