arXiv:2606.05843cs.CLcs.AI2026-06被引 1

发现多模态大模型中专注提取关键视觉信息的特殊注意力头。

Mechanistic Insights into Functional Sparsity in Multimodal LLMs via CoRe Heads

论文配图:Mechanistic Insights into Functional Sparsity in Multimodal LLMs via CoRe Heads
图 1 · 摘自论文原文
  • 通过注意力质量指标识别出专精于跨模态检索的CoRe头。
  • 去掉顶部5%的CoRe头会使多模态推理性能大幅下降。
  • 这类稀疏结构能加速推理且保持高性能,适合模型优化研究者。

尽管多模态大语言模型(MLLMs)在复杂视觉语言任务上表现优异,但其如何从复杂噪声背景中提取与查询相关视觉特征的机制仍不清晰。本文通过引入一种名为检索注意力质量(RAM)的粒度级指标,揭示了MLLMs中一种深层结构特性:跨模态检索的功能稀疏性。我们识别并表征了一类高度专门化的注意力头,称为上下文感知检索(CoRe)头。在多种视觉领域和不同模型规模下,均观察到明确的功能分工:CoRe头作为专用信息提取器,而其他大多数头则将注意力分布于更广泛的上下文区域。因果干预实验进一步证明这些专用头的必要性——仅移除前5%的CoRe头即导致多模态推理性能显著下降,而移除排名较低的头影响甚微。此外,加速实验验证了CoRe头的实用性,表明利用这种局部稀疏性可显著提升推理速度,同时保持稳健的任务表现。本研究揭示了MLLMs中功能稀疏性的结构原则,深化了对机制可解释性的理解,并为未来架构设计与模型优化提供了理论基础。

原文摘要 · Abstract (English)

While Multimodal Large Language Models (MLLMs) demonstrate remarkable proficiency on complex vision-language tasks, the mechanisms by which they extract query-relevant visual features from complex, noisy contexts remain opaque. In this paper, we present an in-depth interpretability study that uncovers a profound structural property within MLLMs: functional sparsity in cross-modal retrieval. Leveraging a token-level metric termed Retrieval Attention Mass (RAM), we identify and characterize a highly specialized subset of attention heads, referred to as Context-aware Retrieval (CoRe) heads. Across diverse visual domains and model scales, we observe a clear functional division: CoRe heads act as dedicated information extractors, while most other heads distribute attention over broader contextual regions. Causal interventions further demonstrate the necessity of these specialized heads. Ablating only the top 5% of CoRe heads causes significant degradation in multimodal reasoning performance, whereas ablating lower-ranked heads has minimal effect. Moreover, acceleration experiments validate the utility of CoRe heads, showing that leveraging this localized sparsity significantly accelerates inference while maintaining robust task performance. Our findings reveal a structural principle of functional sparsity within MLLMs, refining the current understanding of mechanistic interpretability and laying a theoretical foundation that can inspire future architecture design and model optimization.

多模态注意力机制可解释性模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。