发现视觉语言模型依赖虚假关联属性,提出方法提升泛化能力
Black Sheep in the Herd: Playing with Spuriously Correlated Attributes for Vision-Language Recognition
- 识别并过滤导致偏差的虚假相关属性,改善模型决策
- 在11个数据集上提升分布外泛化性能,不牺牲下游任务表现
- 可无缝接入多种高效微调方法,适合关注模型鲁棒性的研究者
少样本适应中的视觉语言模型面临一个困境:在分布内准确率与分布外泛化之间难以平衡。现有研究利用低层视觉属性提升泛化性,但本研究揭示,模型过度依赖少数与类别共现但非本质属性的虚假相关属性。这种偏见导致泛化能力差。为此,本文提出:1)虚假属性探测(SAP),用于识别并过滤此类问题属性,显著提升已有基于属性方法的泛化能力;2)虚假属性屏蔽(SAS),一种即插即用模块,通过削弱这些属性对预测的影响,可无缝集成到多种参数高效微调(PEFT)方法中。实验表明,SAP与SAS在11个数据集和3种泛化任务上均显著提升分布偏移下的准确率,且不损害下游性能,建立新基准。
原文摘要 · Abstract (English)
Few-shot adaptation for Vision-Language Models (VLMs) presents a dilemma: balancing in-distribution accuracy with out-of-distribution generalization. Recent research has utilized low-level concepts such as visual attributes to enhance generalization. However, this study reveals that VLMs overly rely on a small subset of attributes on decision-making, which co-occur with the category but are not inherently part of it, termed spuriously correlated attributes. This biased nature of VLMs results in poor generalization. To address this, 1) we first propose Spurious Attribute Probing (SAP), identifying and filtering out these problematic attributes to significantly enhance the generalization of existing attribute-based methods; 2) We introduce Spurious Attribute Shielding (SAS), a plug-and-play module that mitigates the influence of these attributes on prediction, seamlessly integrating into various Parameter-Efficient Fine-Tuning (PEFT) methods. In experiments, SAP and SAS significantly enhance accuracy on distribution shifts across 11 datasets and 3 generalization tasks without compromising downstream performance, establishing a new state-of-the-art benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。