让视觉模型无视背景干扰,提升图像分类鲁棒性
Birds of a Feather Flock Together: Background-Invariant Representations via Linear Structure in VLMs

- 利用视觉语言模型嵌入空间的线性结构分离前景与背景特征
- 在完全虚假相关条件下实现水鸟数据集90%以上最差组准确率
- 仅用合成数据即可训练,适合真实场景部署
视觉语言模型(如CLIP和SigLIP 2)广泛用于图像分类,但其视觉编码器仍受系统性偏差影响,降低鲁棒性。尤其当前景物体与背景存在显著虚假关联时问题更突出。本文重新审视VLM嵌入空间中高度线性可加的特性,发现该特性支持将场景表示分解为前景与背景成分。基于此,我们提出一种预训练方法,利用合成数据构建对背景不变的表示。据我们所知,该方法在水鸟数据集上实现了首个在100%虚假相关(训练数据中无少数群体样本)条件下的最差组准确率超过90%。此外,该方法展现出优异的从模拟到现实的迁移能力,且无需真实去偏数据,具有实际部署可行性。
原文摘要 · Abstract (English)
Vision-language models (VLMs), such as CLIP and SigLIP 2, are widely used for image classification, yet their vision encoders remain vulnerable to systematic biases that undermine robustness. In particular, correlations between foreground objects and their backgrounds constitute a salient and practically important class of spurious dependencies. In this work, we revisit the well-known property of high linear additivity in VLM embedding spaces and show that it enables a decomposition of scene representations into foreground and background components. Leveraging this insight, we introduce a pre-training approach that exploits this property to construct background-invariant representations using synthetic data. Our method achieves, to our knowledge, the first worst-group accuracy exceeding $90\%$ on Waterbirds under perfect ($100\%$) spurious correlation (i.e., no minority-group examples in the training data). Furthermore, it demonstrates strong sim-to-real transfer and requires no access to real-world debiased data, making it practical for real-world deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。