arXiv:2609.09085cs.CL2026-09

揭示大模型首位置注意力异常的根源,为量化提供新思路

It's Not RoPE that Creates Sinks: The Role of Self-Concentration and Value-Non-Mixing in Attention

论文配图:It's Not RoPE that Creates Sinks: The Role of Self-Concentration and Value-Non-Mixing in Attention
图 1 · 摘自论文原文
  • 发现因果掩码导致注意力自聚焦,引发初始位置异常
  • 注意力输出中值不混合是产生大规模激活的关键原因
  • 适合关注模型内部机制与低比特量化的研究者

大语言模型在序列首位置常出现注意力聚集(AS)和大规模激活(MAs),二者常同时发生,且MAs会阻碍低比特量化。本研究分析了无论首个标记为何都会在初始位置出现AS和MAs的原因。实验表明,因果掩码导致的注意力自聚焦,以及后续注意力输出中的值不混合,是造成该现象的核心因素。该发现为理解大模型注意力层内部动态提供了新的实证支持,有助于未来量化策略的设计,并深化对注意力机制内在机理的认识。

原文摘要 · Abstract (English)

Large Language Models (LLMs) often exhibit "Attention Sink" (AS) and the accompanying "Massive Activations" (MAs) at the initial position of a sequence. These phenomena frequently co-occur, and MAs can pose challenges for low-bit quantization. In this study, we analyze the factors underlying AS and MAs that emerge at the initial position regardless of the token occupying it. Our experiments suggest that self-concentration of attention, resulting from the causal mask, and the subsequent Value-non-mixing in attention outputs contribute to AS and MAs. These findings provide new empirical evidence on the internal dynamics of LLMs, offering insights that may inform future quantization strategies and advance our understanding of the internal mechanisms of attention layers.

注意力机制量化大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。