arXiv:2605.05668cs.AIcs.CV2026-05中稿 · ICML被引 4

发现大模型注意力机制冗余,换随机权重仍表现良好

Large Vision-Language Models Get Lost in Attention

论文配图:Large Vision-Language Models Get Lost in Attention
图 1 · 摘自论文原文
  • 用信息论与几何统一框架分析残差更新机制
  • 注意力层换为高斯噪声权重后多数数据集性能不降反升
  • 揭示当前视觉语言模型严重依赖无效注意力,应重构架构

尽管训练范式快速演进,大型视觉-语言模型(LVLMs)的解码器骨干仍基于残差连接的Transformer架构。因此,解析内部模块的独立作用对理解模型机理和指导架构优化至关重要。以往统计方法虽提供有价值归因见解,但缺乏统一理论基础。为此,我们提出一个基于信息论与几何学的统一框架,量化残差更新的几何与熵特性。应用该框架揭示了根本性功能解耦:注意力作为子空间保持算子,聚焦于重组;前馈网络(FFN)则作为子空间扩展算子,驱动语义创新。进一步实验表明,将学习到的注意力权重替换为预定义值(如高斯噪声),在多数数据集上性能相当甚至更优。这些结果暴露了当前机制中严重的误分配与冗余,表明最先进的LVLM实际上‘迷失在注意力中’,未能有效利用视觉上下文。

原文摘要 · Abstract (English)

Despite the rapid evolution of training paradigms, the decoder backbone of large vision--language models (LVLMs) remains fundamentally rooted in the residual-connection Transformer architecture. Therefore, deciphering the distinct roles of internal modules is critical for understanding model mechanics and guiding architectural optimization. While prior statistical approaches have provided valuable attribution-based insights, they often lack a unified theoretical basis. To bridge this gap, we propose a unified framework grounded in information theory and geometry to quantify the geometric and entropic nature of residual updates. Applying this unified framework reveals a fundamental functional decoupling: Attention acts as a subspace-preserving operator focused on reconfiguration, whereas FFNs serve as subspace-expanding operators driving semantic innovation. Strikingly, further experiments demonstrate that replacing learned attention weights with predefined values (e.g., Gaussian noise) yields comparable or even superior performance across a majority of datasets relative to vanilla models. These results expose severe misallocation and redundancy in current mechanisms, suggesting that state-of-the-art LVLMs effectively ``get lost in attention'' rather than efficiently leveraging visual context.

视觉语言模型注意力机制架构分析信息论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。