arXiv:2505.13946cs.AI2025-05NeurIPS被引 6

通过信息瓶颈原理提升多模态模型在分布外场景下的泛化能力

Visual Instruction Bottleneck Tuning

  • 基于信息瓶颈理论设计可训练的表示学习方法
  • 在45个数据集上显著降低分布偏移下的错误率
  • 适合需要鲁棒性增强的多模态应用开发者

尽管多模态大语言模型(MLLMs)被广泛采用,但在分布外查询下仍会表现下降。现有提升泛化性的方法通常需要更多指令数据或更大模型架构,带来高昂的人力或计算成本。本文从表示学习角度提出新方法,受信息瓶颈(IB)原理启发,推导出MLLM的变分下界并设计实用算法——视觉指令瓶颈调优(Vittle)。理论上证明了Vittle与信息论意义上的模型鲁棒性度量相关。在45个数据集(含30种分布偏移场景)上的开放和闭合问答、物体幻觉检测任务中,多个MLLM均验证:通过学习最小充分表示,Vittle能持续提升模型在分布偏移下的鲁棒性。

原文摘要 · Abstract (English)

Despite widespread adoption, multimodal large language models (MLLMs) suffer performance degradation when encountering unfamiliar queries under distribution shifts. Existing methods to improve MLLM generalization typically require either more instruction data or larger advanced model architectures, both of which incur non-trivial human labor or computational costs. In this work, we take an alternative approach to enhance the generalization and robustness of MLLMs under distribution shifts, from a representation learning perspective. Inspired by information bottleneck (IB) principle, we derive a variational lower bound of the IB for MLLMs and devise a practical implementation, Visual Instruction Bottleneck Tuning (Vittle). We then provide a theoretical justification of Vittle by revealing its connection to an information-theoretic robustness metric of MLLM. Empirical validation of multiple MLLMs on open-ended and closed-form question answering and object hallucination detection tasks over 45 datasets, including 30 shift scenarios, demonstrates that Vittle consistently improves the MLLM's robustness under shifts by pursuing the learning of a minimal sufficient representation.

多模态信息瓶颈鲁棒性表示学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。