arXiv:2604.00503cs.CV2026-04中稿 · CVPR被引 2

PET-DINO通过提示增强训练,提升开放集目标检测的泛化能力。

PET-DINO: Unifying Visual Cues into Grounding DINO with Prompt-Enriched Training

  • 引入视觉提示生成模块与双级提示训练策略
  • 在多个协议下实现竞争力强的零样本检测性能
  • 适合需要快速部署通用目标检测模型的研究者

开放集目标检测(OSOD)可识别超出固定类别的新类别,但面临文本表征与复杂视觉概念对齐困难,以及稀有类别图像-文本配对数据稀缺的问题,导致专业领域或复杂物体上表现不佳。现有视觉提示方法虽部分缓解问题,但常涉及复杂的多模态设计和多阶段优化,延长开发周期。有效数据驱动的OSOD训练策略仍不明确。为此,我们提出PET-DINO,一个支持文本与视觉提示的通用检测器。其提出的对齐友好视觉提示生成(AFVPG)模块基于先进文本提示检测器,克服文本表征引导局限,缩短开发周期。我们设计两种提示增强训练策略:迭代级的同批并行提示(IBP)与训练级的动态记忆驱动提示(DMD),实现多提示路径的并行建模,适应多样化真实场景。大量实验表明,PET-DINO在多种提示检测协议下展现出竞争力的零样本检测能力。其优势源于继承性设计理念与提示增强训练策略,在构建高效通用目标检测器中起关键作用。

原文摘要 · Abstract (English)

Open-Set Object Detection (OSOD) enables recognition of novel categories beyond fixed classes but faces challenges in aligning text representations with complex visual concepts and the scarcity of image-text pairs for rare categories. This results in suboptimal performance in specialized domains or with complex objects. Recent visual-prompted methods partially address these issues but often involve complex multi-modal designs and multi-stage optimizations, prolonging the development cycle. Additionally, effective training strategies for data-driven OSOD models remain largely unexplored. To address these challenges, we propose PET-DINO, a universal detector supporting both text and visual prompts. Our Alignment-Friendly Visual Prompt Generation (AFVPG) module builds upon an advanced text-prompted detector, addressing the limitations of text representation guidance and reducing the development cycle. We introduce two prompt-enriched training strategies: Intra-Batch Parallel Prompting (IBP) at the iteration level and Dynamic Memory-Driven Prompting (DMD) at the overall training level. These strategies enable simultaneous modeling of multiple prompt routes, facilitating parallel alignment with diverse real-world usage scenarios. Comprehensive experiments demonstrate that PET-DINO exhibits competitive zero-shot object detection capabilities across various prompt-based detection protocols. These strengths can be attributed to inheritance-based philosophy and prompt-enriched training strategies, which play a critical role in building an effective generic object detector. Project page: https://fuweifuvtoo.github.io/pet-dino.

目标检测开放集提示学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。