arXiv:2510.06477cs.LGcs.AI2025-10被引 47

发现大模型注意力衰减与压缩谷底实为同一机制的两面。

Attention Sinks and Compression Valleys in LLMs are Two Sides of the Same Coin

  • 通过残差流中的巨大激活解释两种现象的成因。
  • 中层开始的序列首标记激活过强时,两者同步出现。
  • 提出混合-压缩-精炼理论,解释深层计算分工。

注意力衰减和压缩谷底是大语言模型中的两个令人困惑的现象,但以往研究将其孤立看待。本文揭示二者存在惊人关联,根源均在于残差流中大规模激活的形成。理论上证明大规模激活必然导致表征压缩,并给出熵减少的边界。在多个模型(410M至120B参数)上实验验证:当序列起始标记在中层产生极端激活范数时,压缩谷底与注意力衰减同时显现。定向消融实验验证了理论预测。这一统一看法促使我们提出‘混合-压缩-精炼’信息流理论,解释变压器类大模型如何通过大规模激活控制注意力与表征压缩,在深度上组织计算。具体而言,我们认为模型分三阶段运行:(1) 早期广泛混合,(2) 中期压缩计算且混合受限,(3) 晚期选择性精炼。该框架解释为何嵌入任务在中间层表现最佳,而生成任务需全程处理,阐明任务依赖表征差异。

原文摘要 · Abstract (English)

Attention sinks and compression valleys have attracted significant attention as two puzzling phenomena in large language models, but have been studied in isolation. In this work, we present a surprising connection between attention sinks and compression valleys, tracing both to the formation of massive activations in the residual stream. We prove theoretically that massive activations necessarily produce representational compression and establish bounds on the resulting entropy reduction. Through experiments across several models (410M-120B parameters), we confirm that when the beginning-of-sequence token develops extreme activation norms in the middle layers, both compression valleys and attention sinks emerge simultaneously. Targeted ablation studies validate our theoretical predictions. This unified view motivates us to propose the Mix-Compress-Refine theory of information flow, as an attempt to explain how LLMs organize their computation in depth by controlling attention and representational compression via massive activations. Specifically, we posit that Transformer-based LLMs process tokens in three distinct phases: (1) broad mixing in the early layers, (2) compressed computation with limited mixing in the middle layers, and (3) selective refinement in the late layers. Our framework helps explain why embedding tasks perform best at intermediate layers, whereas generation tasks benefit from full-depth processing, clarifying differences in task-dependent representations.

大模型机制注意力机制表征压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。