无需训练,用猜物游戏让模型识别任意物体。
A Training-Free Guess What Vision Language Model from Snippets to Open-Vocabulary Object Detection
- 设计多尺度视觉语言搜索与上下文概念提示机制
- 在多个数据集上达到领先无训练检测性能
- 适合希望快速部署开放词汇检测的开发者
开放词汇目标检测(OVOD)旨在实现对任何物体的检测能力。尽管已有大量大规模预训练工作构建了具备出色零样本能力的基础视觉语言模型,但通常忽略了基于已有预训练模型建立通用理解的必要性。为此,本文提出一种无需训练的“猜物”视觉语言模型(GW-VLM),通过精心设计的多尺度视觉语言搜索(MS-VLS)与上下文概念提示(CCP)机制,实现通用理解范式。该方法让预训练的视觉语言模型(VLM)与大语言模型(LLM)进行‘猜物’互动:MS-VLS利用多尺度视觉-语言软对齐,从无类别目标检测结果中生成片段;CCP则基于这些片段形成概念流,引导LLM理解以完成OVOD。在自然图像和遥感图像数据集(包括COCO val、Pascal VOC、DIOR、NWPU-10)上的大量实验表明,所提GW-VLM在无需任何训练步骤的情况下,显著优于现有最先进方法。
原文摘要 · Abstract (English)
Open-Vocabulary Object Detection (OVOD) aims to develop the capability to detect anything. Although myriads of large-scale pre-training efforts have built versatile foundation models that exhibit impressive zero-shot capabilities to facilitate OVOD, the necessity of creating a universal understanding for any object cognition according to already pretrained foundation models is usually overlooked. Therefore, in this paper, a training-free Guess What Vision Language Model, called GW-VLM, is proposed to form a universal understanding paradigm based on our carefully designed Multi-Scale Visual Language Searching (MS-VLS) coupled with Contextual Concept Prompt (CCP) for OVOD. This approach can engage a pre-trained Vision Language Model (VLM) and a Large Language Model (LLM) in the game of "guess what". Wherein, MS-VLS leverages multi-scale visual-language soft-alignment for VLM to generate snippets from the results of class-agnostic object detection, while CCP can form the concept of flow referring to MS-VLS and then make LLM understand snippets for OVOD. Finally, the extensive experiments are carried out on natural and remote sensing datasets, including COCO val, Pascal VOC, DIOR, and NWPU-10, and the results indicate that our proposed GW-VLM can achieve superior OVOD performance compared to the-state-of-the-art methods without any training step.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。