arXiv:2508.10552cs.CLcs.AI2025-08被引 17

揭示多模态大模型中文字主导现象,提出评估与缓解方法

When Language Overrules: Revealing Text Dominance in Multimodal Large Language Models

  • 首次系统分析图文音时序图等多模态中的文字主导问题
  • 提出MDI和AEI指标,发现文字主导在各模态均显著存在
  • 通过令牌压缩重平衡注意力,使LLaVA-7B的MDI从10.23降至0.86

多模态大语言模型(MLLMs)在多种任务中表现出色,但普遍存在文本主导问题:推理严重依赖文本,而其他模态利用不足。现有研究多归因于数据偏差或架构设计,但未系统考察跨模态影响。本文首次对图像、视频、音频、时间序列和图等多元模态进行系统性分析,提出模态主导指数(MDI)和注意力效率指数(AEI)以量化该不平衡。结果表明,文本主导在所有测试模态中均显著且普遍。深入分析揭示三大成因:非文本模态存在严重令牌冗余导致注意力稀释、融合架构设计影响、任务设定隐含偏好文本输入。进一步提出简单令牌压缩方法,可有效重平衡注意力;应用于LLaVA-7B后,其MDI由10.23降至0.86。本研究为构建更均衡、全面的多模态模型提供基础框架。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities across a diverse range of multimodal tasks. However, these models suffer from a core problem known as text dominance: they depend heavily on text for their inference, while underutilizing other modalities. While prior work has acknowledged this phenomenon in vision-language tasks, often attributing it to data biases or model architectures. In this paper, we conduct the first systematic investigation of text dominance across diverse data modalities, including images, videos, audio, time-series, and graphs. To measure this imbalance, we propose two evaluation metrics: the Modality Dominance Index (MDI) and the Attention Efficiency Index (AEI). Our comprehensive analysis reveals that text dominance is both significant and pervasive across all tested modalities. Our in-depth analysis identifies three underlying causes: attention dilution from severe token redundancy in non-textual modalities, the influence of fusion architecture design, and task formulations that implicitly favor textual inputs. Furthermore, we propose a simple token compression method that effectively rebalances model attention. Applying this method to LLaVA-7B, for instance, drastically reduces its MDI from 10.23 to a well-balanced value of 0.86. Our analysis and methodological framework offer a foundation for the development of more equitable and comprehensive multimodal language models.

多模态文本主导注意力机制模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。