arXiv:2606.01710cs.CVcs.LG2026-06

通过密度感知校准,提升零样本视觉语言模型的鲁棒性

Density-Aware Translation of Spurious Correlations in Zero-Shot VLMs

论文配图:Density-Aware Translation of Spurious Correlations in Zero-Shot VLMs
图 1 · 摘自论文原文
  • 利用嵌入密度重构相似度,抑制无关上下文干扰
  • 在多个数据集上提升最差组与平均准确率
  • 无需微调,适合追求可靠零样本推理的场景

视觉语言模型(如CLIP)虽具备强大的零样本分类能力,但其预测仍受伪相关性影响,即上下文线索压倒语义内容。现有方法多依赖微调或提示工程,易破坏预训练优势或引发幻觉。本文提出密度感知翻译(DAT),基于组参考集的局部几何密度项修正图像-文本相似度。该方法源于CLIP嵌入在特征空间呈现各向异性壳状分布:常见模式聚集于均值附近,稀有模式则被推至外侧,导致对齐不均——伪相关性被放大,而语义有意义但罕见的线索被边缘化。为此,我们引入相对度量,根据嵌入密度重缩放相似度,在稀疏区域抑制过自信得分,同时保留高密度、语义一致的匹配。基准数据集上的实验表明,该方法在最差组和平均准确率上均有稳定提升,验证了密度感知翻译作为简单有效校准机制的潜力。

原文摘要 · Abstract (English)

Vision-Language models (VLMs), such as CLIP, achieve powerful zero-shot classification. However, their predictions remain sensitive to spurious correlations, where contextual cues dominate over semantic content. Earlier solutions typically rely on fine-tuning or prompt engineering, which either undermine the advantages of pre-trained models or are prone to hallucination. In this work, we propose Density-Aware Translation (DAT) that refines image-text similarity scores using a local geometric density term derived from group reference sets. Our approach is motivated by the phenomenon that CLIP embeddings exhibit a modality gap and lie on an anisotropic shell in the feature space: common patterns cluster near the mean, while rare patterns are pushed outward. This geometry creates uneven alignment, where spurious correlations are amplified while semantically meaningful but rare cues are marginalised. To address this, we employ a relative measure to rescale similarities based on embedding density, suppressing overconfident scores in diffuse regions while preserving dense, semantically consistent matches. Experimental results on benchmark datasets demonstrate consistent improvements in worst-group and average accuracy, highlighting density-aware translation as a simple and effective calibration mechanism for reliable zero-shot classification using multimodal models.

零样本学习视觉语言模型伪相关性特征校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。