解决视觉语言模型少样本分类中的分布偏移与过自信问题
Confidence-calibrated covariate shift correction for few-shot classification in Vision-Language Models
- 用费舍尔信息惩罚缓解数据分布偏移
- 通过置信度对齐惩罚降低误分类时的过自信
- 适合需要可靠低资源视觉分类的现实应用
随着视觉语言基础模型在少样本视觉分类任务中的普及,目标数据不足引发的领域泛化问题日益突出。数据稀缺导致采样偏差,并放大模型对数据分布变化的敏感性。尽管在多个领域微调可缓解此问题,但需大量资源和多样化数据源。本文系统分析两大挑战:(1) 预训练分布与目标分布间的协变量偏移;(2) 新数据预测时置信度过高。为此提出统一方法 CalShift,结合费舍尔信息惩罚以缓解协变量偏移,以及置信度对齐惩罚(CMP)以降低误分类样本的过自信。在多种视觉与协变量偏移基准上实验表明,CalShift显著提升模型校准性,预期校准误差(ECE)最高降低5.82%;同时增强鲁棒性,在受协变量偏移影响的挑战性数据集上准确率提升3.5%。结果表明,CalShift是构建真实场景下鲁棒、可靠少样本视觉语言系统的有效策略。
原文摘要 · Abstract (English)
Since the establishment of vision-language foundation models as the new mainstay in low-shot vision classification tasks, the question of domain generalization arising from insufficient target data is assuming more importance. This scarcity challenge induces sampling bias and amplifies model sensitivity to variations and shifts in data distributions. While fine-tuning on multiple domains could mitigate such domain generalization issues, it is resource-intensive and demands diverse data sources. In this work, we systematically analyze two critical challenges: (1) covariate shift between the pre-training distribution and the underspecified target distribution, and (2) confidence misalignment, where predictions on novel data are overconfident. To address both challenges simultaneously, we introduce \textbf{Confidence-Calibrated Covariate Shift Correction (CalShift)} -- a unified approach that combines a Fisher information penalty to mitigate covariate shift and a Confidence Misalignment Penalty (CMP) to reduce overconfidence in misclassified examples. Experimental evaluations across various vision and covariate shift benchmarks demonstrate that CalShift significantly improves model calibration, achieving up to a 5.82\% reduction in Expected Calibration Error (ECE). Furthermore, CalShift enhances robustness, improving accuracy by 3.5\% on challenging datasets impacted by covariate shifts. Our results highlight CalShift as a promising strategy for building robust and reliable low-shot vision-language systems for real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。