arXiv:2502.00662cs.CVcs.CL2025-02被引 5

用图文原型+图像偏差估计,提升少样本分布外检测准确率。

Mitigating the Modality Gap: Few-Shot Out-of-Distribution Detection with Multi-modal Prototypes and Image Bias Estimation

  • 引入图文双原型,缓解图像与文本模态差异带来的误判。
  • 在多个数据集上实现比现有方法更高的检测准确率,无需额外训练。
  • 适合关注少样本分布外检测的视觉语言模型研究者。

基于视觉语言模型(VLM)的分布外(OOD)检测方法通常依赖输入图像与已知类别(ID)文本原型之间的相似度得分。然而,图像与文本间的模态差距常导致高误报率,因为分布外样本可能与ID文本原型表现出高相似性。为缓解该问题,本文提出同时引入ID图像原型和ID文本原型。理论分析与实验证明,该方法可在无需额外训练的情况下提升VLM-based OOD检测性能。为进一步缩小模态差距,我们设计了一种新颖的少样本微调框架SUPREME,包含偏置提示生成(BPG)和图像-文本一致性(ITC)模块:BPG通过基于高斯估计的图像域偏置条件化文本原型,增强图文融合与泛化能力;ITC通过最小化模态内与跨模态距离来降低模态差距。此外,基于理论与实证发现,我们提出一种新型OOD评分 $S_{\textit{GMP}}$,融合单模态与跨模态相似性。大量实验表明,SUPREME在多个数据集上持续优于现有VLM-based OOD检测方法。

原文摘要 · Abstract (English)

Existing vision-language model (VLM)-based methods for out-of-distribution (OOD) detection typically rely on similarity scores between input images and in-distribution (ID) text prototypes. However, the modality gap between image and text often results in high false positive rates, as OOD samples can exhibit high similarity to ID text prototypes. To mitigate the impact of this modality gap, we propose incorporating ID image prototypes along with ID text prototypes. We present theoretical analysis and empirical evidence indicating that this approach enhances VLM-based OOD detection performance without any additional training. To further reduce the gap between image and text, we introduce a novel few-shot tuning framework, SUPREME, comprising biased prompts generation (BPG) and image-text consistency (ITC) modules. BPG enhances image-text fusion and improves generalization by conditioning ID text prototypes on the Gaussian-based estimated image domain bias; ITC reduces the modality gap by minimizing intra- and inter-modal distances. Moreover, inspired by our theoretical and empirical findings, we introduce a novel OOD score $S_{\textit{GMP}}$, leveraging uni- and cross-modal similarities. Finally, we present extensive experiments to demonstrate that SUPREME consistently outperforms existing VLM-based OOD detection methods.

分布外检测视觉语言模型少样本学习模态对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。