提出新框架解释为何模型有时忽略视觉信息,还能自适应调整注意力。
Taming Visual Neglect: A Variational Information Bottleneck Framework for Adaptive Attention in Multimodal In-Context Learning

- 基于信息瓶颈理论,用跨模态增益量化视觉信息价值。
- 实验证明视觉非冗余时多模态学习比纯文本提升4.7%准确率。
- 适合研究多模态模型决策机制或优化注意力分配的研究者。
大型多模态模型具备强大的上下文学习能力,但何时及为何视觉上下文能提升效果仍不清晰。实证发现模型有时有效利用视觉示例,有时却完全忽略。本文提出VIB-ICL框架,基于信息瓶颈原理解决这一矛盾。引入跨模态信息增益(CMIG),量化视觉信息相对于文本信息的额外互信息。推导出泛化界,证明当视觉信息非冗余时,多模态上下文学习的过拟合风险低于纯文本学习,理论上可超越文本仅有的学习。进一步证明,当视觉信息冗余时,忽略视觉是信息瓶颈最优解,导出闭式注意力重分配原则。在VIB-ICL算法中,通过变分界估计CMIG并动态调整注意力。五个基准测试显示,准确率最高提升4.7%,所需演示减少35%,验证了理论预测。
原文摘要 · Abstract (English)
Large vision-language models exhibit strong in-context learning (ICL) capabilities, yet when and why visual context helps multimodal ICL remains poorly understood. Empirical studies show a puzzling dichotomy: models sometimes effectively leverage visual demonstrations, yet often neglect them entirely. We propose VIB-ICL, an information-theoretic framework that resolves this dichotomy through the Information Bottleneck principle. We introduce the Cross-Modal Information Gain (CMIG), which quantifies the additional mutual information that visual context provides about the target beyond textual context. We derive a generalization bound showing that multimodal ICL's excess risk over text-only ICL is governed by the CMIG, proving that multimodal ICL provably outperforms text-only ICL when visual information is non-redundant. We further prove that visual context neglect, often viewed as a failure mode, is the Information Bottleneck-optimal solution when visual information is redundant, yielding a closed-form Attention Reallocation Principle that prescribes how visual attention weights should be adaptively adjusted. We instantiate this principle in the VIB-ICL algorithm, which estimates CMIG via variational bounds and dynamically reallocates attention. Experiments on five benchmarks demonstrate consistent improvements of up to 4.7\% accuracy gains and 35\% reduction in required demonstrations, validating our theoretical predictions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。