用视觉提示和文本增强,让模型零样本识别物体级异常
VisTa: Visual-contextual and Text-augmented Zero-shot Object-level OOD Detection
- 用视觉提示保留上下文信息,提升定位能力
- 在COCO、OpenImages等数据集上表现优于现有方法
- 适合部署在无法访问训练数据的检测系统中
随着目标检测器越来越多地以黑箱云服务或预训练模型形式部署,且无法访问原始训练数据,零样本物体级分布外(OOD)检测成为关键挑战。尽管现有方法利用CLIP等预训练视觉语言模型在图像级OOD检测上取得成功,但直接应用于物体级检测时会丢失上下文信息并依赖图像级对齐。为此,本文提出一种新方法,通过视觉提示与文本增强的分布内(ID)空间构建,适配CLIP实现零样本物体级OOD检测。该方法有效保留了关键上下文信息,显著提升了区分分布内与分布外物体的能力,在多个基准测试中均达到竞争性性能。
原文摘要 · Abstract (English)
As object detectors are increasingly deployed as black-box cloud services or pre-trained models with restricted access to the original training data, the challenge of zero-shot object-level out-of-distribution (OOD) detection arises. This task becomes crucial in ensuring the reliability of detectors in open-world settings. While existing methods have demonstrated success in image-level OOD detection using pre-trained vision-language models like CLIP, directly applying such models to object-level OOD detection presents challenges due to the loss of contextual information and reliance on image-level alignment. To tackle these challenges, we introduce a new method that leverages visual prompts and text-augmented in-distribution (ID) space construction to adapt CLIP for zero-shot object-level OOD detection. Our method preserves critical contextual information and improves the ability to differentiate between ID and OOD objects, achieving competitive performance across different benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。