发现softmax注意力的隐藏结构规律,揭示其内在数学特性与训练监控新方法。
On the Invariants of Softmax Attention

- 提出能量场概念,揭示注意力对数的行中心不变性。
- 发现每行和为零、秩受头维度限制等数学约束,且在多种模型中成立。
- 关键矩阵的非相干性可作训练监控指标,适用于各类自回归语言模型。
Softmax注意力将每个查询-键交互映射为概率分布,但其底层结构仍不明确。本文定义了‘能量场’(即行中心化的注意力对数),发现其在不同模型、架构和输入下具有不变性。两类不变性浮现:机制级不变性源于softmax注意力的代数结构,包括每行和为零、由头维度决定的秩上界,以及由此产生的谱特征;模型级规律虽非机制必需,但在测试的所有自回归语言模型中均成立,涵盖多个架构家族。能量场的方差分布在多个键位置,而非集中于少数位置,这种分散性源于我们称之为‘关键矩阵非相干性’的性质。这些不变性具有实际意义:秩上界使能量场受限于低维子空间,非相干性则提供了每头的训练监控手段。所有结果在多种上下文长度和文本输入下均得到验证。
原文摘要 · Abstract (English)
Softmax attention maps every query--key interaction into a probability distribution, but the underlying structure remains largely unexplored. We define the \emph{energy field}, the row-centered attention logit, and show that it exhibits invariant properties across models, architectures, and inputs. Two classes of invariants emerge. \emph{Mechanism-level} invariants follow from the algebraic structure of softmax attention. They include a per-row zero-sum constraint, a rank bound determined by the head dimension, and spectral signatures that follow from them. \emph{Model-level} regularities are not required by the mechanism, yet hold in every autoregressive language model we test, spanning several architecture families. The energy field distributes its variance over key positions without concentrating at a few. This delocalization traces to a property of the key matrix we call \emph{key incoherence}. These invariants have practical consequences. The rank bound confines the energy field to a low-dimensional subspace. Key incoherence yields a per-head training monitor. All results are verified at multiple context lengths and input texts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。