通过信息瓶颈原理改进多模态对齐,让不同模态表示更一致。
Aligning Multimodal Representations through an Information Bottleneck
- 基于信息瓶颈理论,指出对比损失无法去除模态特有信息。
- 引入可学习正则项,显著提升图像与文本表示的对齐程度。
- 适合关注多模态对齐机制或希望提升模型泛化能力的研究者。
对比损失广泛用于多模态表征学习,但实证发现其难以有效学习对齐的表征空间。本文认为该现象源于表征空间中存在模态特有信息。尽管许多常用对比损失旨在最大化双模态表征间的互信息,却未设计用于消除模态特有信息。我们从信息瓶颈原理出发,对该问题提供理论解释,并在受控实验中分析不同超参数对这一现象的影响。最后,提出一种基于变分近似的损失正则项,旨在增强表征对齐。通过一系列受控实验和真实应用场景验证,该正则项能显著提升模型性能。
原文摘要 · Abstract (English)
Contrastive losses have been extensively used as a tool for multimodal representation learning. However, it has been empirically observed that their use is not effective to learn an aligned representation space. In this paper, we argue that this phenomenon is caused by the presence of modality-specific information in the representation space. Although some of the most widely used contrastive losses maximize the mutual information between representations of both modalities, they are not designed to remove the modality-specific information. We give a theoretical description of this problem through the lens of the Information Bottleneck Principle. We also empirically analyze how different hyperparameters affect the emergence of this phenomenon in a controlled experimental setup. Finally, we propose a regularization term in the loss function that is derived by means of a variational approximation and aims to increase the representational alignment. We analyze in a set of controlled experiments and real-world applications the advantages of including this regularization term.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。