arXiv:2604.14684cs.CV2026-04被引 1

提出DETR-ViP,让视觉提示更精准区分物体类别

DETR-ViP: Detection Transformer with Robust Discriminative Visual Prompts

论文配图:DETR-ViP: Detection Transformer with Robust Discriminative Visual Prompts
图 1 · 摘自论文原文
  • 通过全局提示融合与图文关系蒸馏,增强视觉提示的区分能力
  • 在COCO、LVIS等数据集上显著提升稀有类别检测效果
  • 适合需要灵活定义目标类别的开放词汇检测场景

视觉提示目标检测支持交互式、灵活的目标类别定义,推动开放词汇检测发展。由于视觉提示直接来自图像特征,其对罕见类别的识别性能通常优于文本提示。然而,视觉提示检测研究长期被忽视,常被视为训练文本提示检测器的副产品,限制了其发展。本文揭示其性能不佳的根本原因在于视觉提示缺乏全局区分性。为此,提出DETR-ViP框架,通过基础图像-文本对比学习,引入全局提示整合与图文提示关系蒸馏,学习更具区分性的提示表示,并采用选择性融合策略确保检测稳定可靠。在COCO、LVIS、ODinW和Roboflow100上的大量实验表明,DETR-ViP在视觉提示检测任务上显著优于现有先进方法。一系列消融实验与分析进一步验证了所提改进的有效性,揭示了视觉提示检测能力提升的内在机制。

原文摘要 · Abstract (English)

Visual prompted object detection enables interactive and flexible definition of target categories, thereby facilitating open-vocabulary detection. Since visual prompts are derived directly from image features, they often outperform text prompts in recognizing rare categories. Nevertheless, research on visual prompted detection has been largely overlooked, and it is typically treated as a byproduct of training text prompted detectors, which hinders its development. To fully unlock the potential of visual-prompted detection, we investigate the reasons why its performance is suboptimal and reveal that the underlying issue lies in the absence of global discriminability in visual prompts. Motivated by these observations, we propose DETR-ViP, a robust object detection framework that yields class-distinguishable visual prompts. On top of basic image-text contrastive learning, DETR-ViP incorporates global prompt integration and visual-textual prompt relation distillation to learn more discriminative prompt representations. In addition, DETR-ViP employs a selective fusion strategy that ensures stable and robust detection. Extensive experiments on COCO, LVIS, ODinW, and Roboflow100 demonstrate that DETR-ViP achieves substantially higher performance in visual prompt detection compared to other state-of-the-art counterparts. A series of ablation studies and analyses further validate the effectiveness of the proposed improvements and shed light on the underlying reasons for the enhanced detection capability of visual prompts.

目标检测视觉提示开放词汇DETR

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。