提出新分类框架,用图文双模态视角重看CLIP类模型的分布外检测。
Recent Advances in Out-of-Distribution Detection with CLIP-Like Models: A Survey
- 按图像与文本模态的利用方式分四类,突破传统单模态局限
- 支持无训练与需训练两种策略,覆盖未知图像/文本场景
- 指出现有挑战,推动跨领域融合与理论研究
分布外检测(OOD)是现实应用中识别测试阶段与训练数据分布不同的样本的关键任务。近年来,以CLIP为代表的视觉语言模型(VLM)推动了该领域的变革,从传统的单模态图像检测转向多模态图像-文本检测。现有分类方法(如少样本或零样本类型)仍依赖于可用的分布内(ID)图像,遵循单模态范式。为更好契合CLIP的跨模态特性,本文提出基于图像与文本模态的新型分类框架:根据对分布外(OOD)数据的视觉与文本信息使用方式,将方法分为四类——已知或未知的OOD图像,以及已知或未知的OOD文本(即可学习向量或类别名),并结合两种训练策略(免训练或需训练)。此外,本文探讨了当前CLIP类模型在分布外检测中的开放问题,展望未来研究方向,包括跨域整合、实际应用及理论理解。
原文摘要 · Abstract (English)
Out-of-distribution detection (OOD) is a pivotal task for real-world applications that trains models to identify samples that are distributionally different from the in-distribution (ID) data during testing. Recent advances in AI, particularly Vision-Language Models (VLMs) like CLIP, have revolutionized OOD detection by shifting from traditional unimodal image detectors to multimodal image-text detectors. This shift has inspired extensive research; however, existing categorization schemes (e.g., few- or zero-shot types) still rely solely on the availability of ID images, adhering to a unimodal paradigm. To better align with CLIP's cross-modal nature, we propose a new categorization framework rooted in both image and text modalities. Specifically, we categorize existing methods based on how visual and textual information of OOD data is utilized within image + text modalities, and further divide them into four groups: OOD Images (i.e., outliers) Seen or Unseen, and OOD Texts (i.e., learnable vectors or class names) Known or Unknown, across two training strategies (i.e., train-free or training-required). More importantly, we discuss open problems in CLIP-like OOD detection and highlight promising directions for future research, including cross-domain integration, practical applications, and theoretical understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。