分离视觉语言中的不变与虚假特征,提升模型对新类别和分布外数据的泛化能力
DiMPLe -- Disentangled Multi-Modal Prompt Learning: Enhancing Out-Of-Distribution Alignment with Invariant and Spurious Feature Separation
- 通过多模态提示学习,分离图像与文本中的不变与虚假特征
- 在11个数据集上平均提升基类准确率15.27点,新类别准确率44.31点
- 适合关注跨模态泛化与分布外鲁棒性的研究者使用
我们提出DiMPLe(解耦多模态提示学习),一种在多模态学习中解耦视觉与语言模态下不变与虚假特征的新方法。视觉数据中的虚假相关性常导致分布外(OOD)性能下降。不同于仅关注图像特征的先前方法,DiMPLe在模态内与模态间同时解耦特征,保持一致对齐,从而实现对新类别更好的泛化能力与对分布偏移的鲁棒性。该方法结合三项关键目标:(1) 不变与虚假特征间互信息最小化,(2) 虚假特征正则化,(3) 不变特征上的对比学习。大量实验表明,相较于CoOp-OOD,DiMPLe在11个多样化数据集上的平均表现更优,基类准确率提升15.27点,新类别准确率提升44.31点。
原文摘要 · Abstract (English)
We introduce DiMPLe (Disentangled Multi-Modal Prompt Learning), a novel approach to disentangle invariant and spurious features across vision and language modalities in multi-modal learning. Spurious correlations in visual data often hinder out-of-distribution (OOD) performance. Unlike prior methods focusing solely on image features, DiMPLe disentangles features within and across modalities while maintaining consistent alignment, enabling better generalization to novel classes and robustness to distribution shifts. Our method combines three key objectives: (1) mutual information minimization between invariant and spurious features, (2) spurious feature regularization, and (3) contrastive learning on invariant features. Extensive experiments demonstrate DiMPLe demonstrates superior performance compared to CoOp-OOD, when averaged across 11 diverse datasets, and achieves absolute gains of 15.27 in base class accuracy and 44.31 in novel class accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。