分析视觉语言模型如何零样本检测异常数据,揭示其优势与敏感点
An Empirical Analysis of VLM-based OOD Detection: Mechanisms, Advantages, and Sensitivity
- 通过嵌入空间特性解析VLM实现零样本异常检测的机制
- 相比单模态方法,语义新颖性使VLM检测效果更优
- 对提示词变化敏感,但对图像噪声有较强鲁棒性
视觉语言模型(如CLIP)展现出卓越的零样本异常检测能力,对构建可靠AI系统至关重要。然而,当前研究对(1)其高效原因、(2)相比单模态方法的优势、(3)行为鲁棒性等关键问题仍缺乏系统理解。本文通过使用分布内(ID)和分布外(OOD)提示,对基于VLM的异常检测进行系统性实证分析。首先,系统刻画并形式化了VLM嵌入空间中促进零样本异常检测的关键操作属性;其次,实证量化了此类模型相较于经典单模态方法的优越性,归因于其利用丰富语义新颖性的能力;最后,发现其鲁棒性存在显著且此前未被充分关注的不对称性:虽对常见图像噪声具有韧性,却对提示词表述极为敏感。研究成果为深入理解VLM异常检测的内在优势与关键缺陷提供了结构化认知,为未来更鲁棒、可靠的系统设计提供实证指导。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs), such as CLIP, have demonstrated remarkable zero-shot out-of-distribution (OOD) detection capabilities, vital for reliable AI systems. Despite this promising capability, a comprehensive understanding of (1) why they work so effectively, (2) what advantages do they have over single-modal methods, and (3) how is their behavioral robustness -- remains notably incomplete within the research community. This paper presents a systematic empirical analysis of VLM-based OOD detection using in-distribution (ID) and OOD prompts. (1) Mechanisms: We systematically characterize and formalize key operational properties within the VLM embedding space that facilitate zero-shot OOD detection. (2) Advantages: We empirically quantify the superiority of these models over established single-modal approaches, attributing this distinct advantage to the VLM's capacity to leverage rich semantic novelty. (3) Sensitivity: We uncovers a significant and previously under-explored asymmetry in their robustness profile: while exhibiting resilience to common image noise, these VLM-based methods are highly sensitive to prompt phrasing. Our findings contribute a more structured understanding of the strengths and critical vulnerabilities inherent in VLM-based OOD detection, offering crucial, empirically-grounded guidance for developing more robust and reliable future designs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。