用信息瓶颈理论解释CLIP为何泛化好,并提升其性能
CIBR: Cross-modal Information Bottleneck Regularization for Robust CLIP Generalization
- 从信息瓶颈视角重释CLIP对比学习目标
- 在7个图像数据集上零样本分类均提升,检索任务也更优
- 适合研究跨模态学习机制或想改进CLIP的学者
对比语言-图像预训练(CLIP)在零样本图像分类和图文检索等跨模态任务中取得显著成功,通过有效对齐视觉与文本表示实现。然而,其强大泛化能力的理论基础仍不明确。本文提出跨模态信息瓶颈(CIB)框架,将CLIP的对比学习目标视为隐式的信息瓶颈优化:模型最大化跨模态共享信息,同时丢弃模态特异性冗余,从而保留跨模态的语义对齐。基于此,我们提出跨模态信息瓶颈正则化(CIBR)方法,在训练中显式施加该原则,引入惩罚项以抑制模态特异性冗余,增强图像与文本特征间的语义对齐。我们在多个视觉-语言基准上验证了CIBR,包括七个不同图像数据集上的零样本分类以及MSCOCO和Flickr30K上的图文检索任务。结果表明,相较标准CLIP,CIBR表现持续提升。这些发现首次通过信息瓶颈视角揭示了CLIP泛化能力的理论机制,同时也带来实际性能改进,为未来跨模态表示学习提供指导。
原文摘要 · Abstract (English)
Contrastive Language-Image Pretraining (CLIP) has achieved remarkable success in cross-modal tasks such as zero-shot image classification and text-image retrieval by effectively aligning visual and textual representations. However, the theoretical foundations underlying CLIP's strong generalization remain unclear. In this work, we address this gap by proposing the Cross-modal Information Bottleneck (CIB) framework. CIB offers a principled interpretation of CLIP's contrastive learning objective as an implicit Information Bottleneck optimization. Under this view, the model maximizes shared cross-modal information while discarding modality-specific redundancies, thereby preserving essential semantic alignment across modalities. Building on this insight, we introduce a Cross-modal Information Bottleneck Regularization (CIBR) method that explicitly enforces these IB principles during training. CIBR introduces a penalty term to discourage modality-specific redundancy, thereby enhancing semantic alignment between image and text features. We validate CIBR on extensive vision-language benchmarks, including zero-shot classification across seven diverse image datasets and text-image retrieval on MSCOCO and Flickr30K. The results show consistent performance gains over standard CLIP. These findings provide the first theoretical understanding of CLIP's generalization through the IB lens. They also demonstrate practical improvements, offering guidance for future cross-modal representation learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。