arXiv:2503.20492cs.CVcs.AI2025-03被引 2

用视觉语言模型实现高效通用的少样本误分类检测

Towards Efficient and General-Purpose Few-Shot Misclassification Detection for Vision-Language Models

  • 基于提示学习构建少样本误分类检测框架,无需从头训练
  • 通过自适应伪样本生成和负损失,降低模型过自信问题
  • 在多个数据集上验证泛化能力,适合动态变化场景

可靠的分类器预测对高安全性和动态环境中的部署至关重要。然而,现代神经网络在错误分类时往往表现出过度自信,凸显了置信度估计以检测错误的必要性。尽管现有方法在小规模数据集上取得进展,但均需从头训练,缺乏高效有效的误分类检测(MisD)方法,限制了其在大规模、不断变化数据集上的应用。本文提出利用视觉语言模型(VLM)的文本信息,构建一个高效且通用的误分类检测框架。通过提示学习,我们设计了FSMisD,避免从头训练,提升调优效率。为增强检测能力,引入自适应伪样本生成与新型负损失,通过将类别提示远离伪特征缓解过自信问题。在多种提示学习方法及存在领域偏移的数据集上进行综合实验,结果表明该方法在效果、效率和泛化性上均有显著且一致的提升。

原文摘要 · Abstract (English)

Reliable prediction by classifiers is crucial for their deployment in high security and dynamically changing situations. However, modern neural networks often exhibit overconfidence for misclassified predictions, highlighting the need for confidence estimation to detect errors. Despite the achievements obtained by existing methods on small-scale datasets, they all require training from scratch and there are no efficient and effective misclassification detection (MisD) methods, hindering practical application towards large-scale and ever-changing datasets. In this paper, we pave the way to exploit vision language model (VLM) leveraging text information to establish an efficient and general-purpose misclassification detection framework. By harnessing the power of VLM, we construct FSMisD, a Few-Shot prompt learning framework for MisD to refrain from training from scratch and therefore improve tuning efficiency. To enhance misclassification detection ability, we use adaptive pseudo sample generation and a novel negative loss to mitigate the issue of overconfidence by pushing category prompts away from pseudo features. We conduct comprehensive experiments with prompt learning methods and validate the generalization ability across various datasets with domain shift. Significant and consistent improvement demonstrates the effectiveness, efficiency and generalizability of our approach.

少样本学习误分类检测视觉语言模型置信度估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。