arXiv:2505.18657cs.AI2025-05被引 38

MLLMs严重依赖语言模态,忽视视觉信息,本文揭示其根源并提出改进路径。

MLLMs are Deeply Affected by Modality Bias

  • 分析三类导致模态偏见的根源:数据特性、模型架构与训练目标
  • 实验证明语言数据更紧凑,视觉数据更冗余,加剧学习失衡
  • 适合关注多模态模型公平性与鲁棒性的研究者阅读

多模态大语言模型(MLLMs)在整合文本与图像等异构模态方面取得进展,但严重受制于模态偏见,常过度依赖语言而忽视视觉输入。本文指出,当前模态偏见在各类任务中普遍存在。我们系统诊断了偏见成因:1. 数据特性——语言数据紧凑抽象,视觉数据冗余复杂,导致学习动态失衡;2. 骨干网络能力不均——预训练语言模型主导,使视觉信息被忽略;3. 训练目标缺陷——现有目标未能促进跨模态均衡对齐,引发语言导向的捷径学习。实验验证了这些因素的影响,强调需采用平衡的训练策略与架构设计以更好融合多模态。本文呼吁跨学科合作,推动更稳健、可泛化的多模态系统发展,助力通用人工智能进步。

原文摘要 · Abstract (English)

Recent advances in Multimodal Large Language Models (MLLMs) have shown promising results in integrating diverse modalities such as texts and images. MLLMs are heavily influenced by modality bias, often relying on language while under-utilizing other modalities like visual inputs. This position paper argues that MLLMs are deeply affected by modality bias. Firstly, we diagnose the current state of modality bias, highlighting its manifestations across various tasks. Secondly, we propose a systematic research road-map related to modality bias in MLLMs. Thirdly, we identify key factors of modality bias in MLLMs and offer actionable suggestions for future research to mitigate it. To substantiate these findings, we conduct experiments that demonstrate the influence of each factor: 1. Data Characteristics: Language data is compact and abstract, while visual data is redundant and complex, creating an inherent imbalance in learning dynamics. 2. Imbalanced Backbone Capabilities: The dominance of pretrained language models in MLLMs leads to overreliance on language and neglect of visual information. 3. Training Objectives: Current objectives often fail to promote balanced cross-modal alignment, resulting in shortcut learning biased toward language. These findings highlight the need for balanced training strategies and model architectures to better integrate multiple modalities in MLLMs. We call for interdisciplinary efforts to tackle these challenges and drive innovation in MLLM research. Our work provides a fresh perspective on modality bias in MLLMs and offers insights for developing more robust and generalizable multimodal systems-advancing progress toward Artificial General Intelligence.

多模态模型偏见语言依赖训练策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。