提升视觉语言模型对齐效果,让图文理解更平衡
VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization
- 通过最大化跨模态互信息,显式优化图文对齐
- 在多个基准上超越基线,尤其在长文本下表现更优
- 无需新增参数或数据,可直接增强现有模型
当前多模态大语言模型在模态对齐方面存在严重挑战,常过度依赖文本信息而忽视视觉模态。本文对广泛使用的交叉熵损失进行了信息论分析,发现其隐含的对齐目标存在固有缺陷,导致文本序列越长,跨模态对齐性能越差,影响多模态信息融合。为此,我们提出基于理论洞察的VISTA方法,引入显式对齐目标以最大化跨模态互信息,有效防止视觉对齐退化。VISTA无需额外可训练模块或训练数据,即可显著提升现有MLLM的视觉理解能力,在超过十余个基准数据集(如VQAv2、MMStar、MME)上表现优异,为多模态对齐研究开辟新方向。
原文摘要 · Abstract (English)
Current multimodal large language models (MLLMs) face a critical challenge in modality alignment, often exhibiting a bias towards textual information at the expense of other modalities like vision. This paper conducts a systematic information-theoretic analysis of the widely used cross-entropy loss in MLLMs, uncovering its implicit alignment objective. Our theoretical investigation reveals that this implicit objective has inherent limitations, leading to a degradation of cross-modal alignment as text sequence length increases, thereby hindering effective multimodal information fusion. To overcome these drawbacks, we propose Vision-Text Alignment (VISTA), a novel approach guided by our theoretical insights. VISTA introduces an explicit alignment objective designed to maximize cross-modal mutual information, preventing the degradation of visual alignment. Notably, VISTA enhances the visual understanding capabilities of existing MLLMs without requiring any additional trainable modules or extra training data, making it both efficient and practical. Our method significantly outperforms baseline models across more than a dozen benchmark datasets, including VQAv2, MMStar, and MME, paving the way for new directions in MLLM modal alignment research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。