arXiv:2605.26661cs.CVcs.AI2026-05

用视觉特征学习类别原型,解决图文模型的模态差异问题

Respecting Modality Gap in Post-hoc Out-of-distribution Detection with Pre-trained Vision-Language Models

  • 在测试时用无标签数据流和软预测在线学习视觉原型
  • 相比现有方法,跨多个设置实现最新性能
  • 适合需要可靠异常检测的视觉语言应用

分布外(OOD)检测已成为提升机器学习模型可靠性的重要技术,用于识别未知类别的意外输入。预训练视觉语言模型(VLMs)的进展使无需分布内(ID)训练数据即可实现零样本OOD检测成为可能;当前方法通常将类别名称的文本嵌入作为类别原型。本文通过理论分析挑战了这一广泛采用的文本作为原型范式,指出现成文本原型与最优视觉原型存在固有模态差距,仅靠提示工程无法消除。为缓解该差距,在后处理约束下,本文提出一种在线伪监督框架,利用无标签测试数据流和预训练VLM的软预测,直接在视觉特征空间中学习类别原型。我们提供了在线优化过程收敛性的理论保证。大量实验表明,该方法在多种OOD检测设置下均达到新基准性能。

原文摘要 · Abstract (English)

Out-of-distribution (OOD) detection has emerged as a popular technique to enhance the reliability of machine learning models by identifying unexpected inputs from unknown classes. Recent progress in pre-trained vision-language models (VLMs) has enabled zero-shot OOD detection without access to in-distribution (ID) training data; in this setting, existing methods commonly treat text embeddings of class names as class prototypes. In this paper, we challenge the widely adopted text-as-prototype paradigm by theoretically showing that off-the-shelf textual prototypes are generally misaligned with the optimal visual prototypes, yielding an intrinsic modality gap that cannot be eliminated by prompt engineering alone. To mitigate this gap under the post-hoc constraint, this paper presents an online pseudo-supervised framework that directly learns class prototypes in the visual feature space using unlabeled test-time data streams and soft predictions from the pre-trained VLMs. We provide theoretical guarantees for the convergence of the online optimization procedure. Extensive experiments empirically demonstrate that our method achieves a new state of the art across a variety of OOD detection setups.

OOD检测视觉语言模型模态对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。