arXiv:2412.10719cs.CVcs.AI2024-12AAAI被引 5

用几张图片做提示,实现无需人工干预的开放集视觉感知。

Just a Few Glances: Open-Set Visual Perception with Image Prompt Paradigm

  • 用少量图像实例作为提示,自动编码融合完成检测分割
  • 在公开数据集上性能媲美文本/视觉提示方法,专用数据集上显著领先
  • 适合自动化系统中快速识别新类别,尤其适用于专业领域

为突破预训练模型对固定类别的限制,开放集目标检测(OSOD)和开放集分割(OSS)受到广泛关注。受大语言模型启发,主流方法多采用文本提示,取得显著效果。遵循SAM范式,部分研究使用点、框、掩码等视觉提示。然而,这两种范式存在固有缺陷:文本难以准确描述专业类别特征;现有视觉提示依赖多轮人工交互,难以用于全自动流程。为此,本文提出一种新型提示范式——图像提示范式,实现无需多轮人工干预的开放集感知。该范式仅需少量图像实例作为提示,提出名为MI Grounding的新框架,在单阶段推理中自动编码、选择与融合高质量图像提示。在多个公开数据集上的实验表明,MI Grounding在OSOD与OSS基准上性能优于或相当文本提示与视觉提示方法;尤其在自建的专业级ADR50K数据集上表现显著更优。

原文摘要 · Abstract (English)

To break through the limitations of pre-training models on fixed categories, Open-Set Object Detection (OSOD) and Open-Set Segmentation (OSS) have attracted a surge of interest from researchers. Inspired by large language models, mainstream OSOD and OSS methods generally utilize text as a prompt, achieving remarkable performance. Following SAM paradigm, some researchers use visual prompts, such as points, boxes, and masks that cover detection or segmentation targets. Despite these two prompt paradigms exhibit excellent performance, they also reveal inherent limitations. On the one hand, it is difficult to accurately describe characteristics of specialized category using textual description. On the other hand, existing visual prompt paradigms heavily rely on multi-round human interaction, which hinders them being applied to fully automated pipeline. To address the above issues, we propose a novel prompt paradigm in OSOD and OSS, that is, \textbf{Image Prompt Paradigm}. This brand new prompt paradigm enables to detect or segment specialized categories without multi-round human intervention. To achieve this goal, the proposed image prompt paradigm uses just a few image instances as prompts, and we propose a novel framework named \textbf{MI Grounding} for this new paradigm. In this framework, high-quality image prompts are automatically encoded, selected and fused, achieving the single-stage and non-interactive inference. We conduct extensive experiments on public datasets, showing that MI Grounding achieves competitive performance on OSOD and OSS benchmarks compared to text prompt paradigm methods and visual prompt paradigm methods. Moreover, MI Grounding can greatly outperform existing method on our constructed specialized ADR50K dataset.

开放集检测图像提示自动化感知专用识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。