揭示多头注意力动态机制,解释为何聚类稳定且不同头协同加速聚集。
Gradient Flow Structure and Quantitative Dynamics of Multi-Head Self-Attention

- 将多头注意力建模为球面上的梯度流,分析能量函数随时间变化规律。
- 发现径向投影项是阻碍单头单调性的关键,提出保证单调性的充分条件。
- 在简化模型中推导出聚类临界温度公式,证明头间异质性提升聚集速度。
Transformer 自注意力可被解释为在单位球面上的梯度流,其中标记通过 softmax 相互作用势能演化并趋向形成簇。尽管先前工作已确立单头注意力的聚类行为,但由于多头间的几何干扰,导致标准单调性论证失效,多头设置仍不清晰。本文建立多头自注意力动态的理论框架,解决了若干开放问题。在评分矩阵满足特定条件时,证明自然的多头能量泛函在平坦与球面动力学下均非递减。识别出阻碍每头单调性的关键障碍为径向阴影项——即各头输出在标记方向上的投影,即使在正交假设下仍存在。提出确保单调性的充分条件,并证明对近似正交性的鲁棒性。在标量头简化模型中,针对等角标记配置,推导出决定聚类行为的临界逆温度闭式表达式,并表明异质头具有超加性聚类速率。在此模型下,还证明了线性化动力学中 ReLU 与 softmax 注意力的聚类时间分离。最终,建立熵产生恒等式,证明注意力熵随聚类进展单调增加至平衡态。结果提供多头注意力动态的统一视角,阐明聚类与稳定性机制。
原文摘要 · Abstract (English)
Transformer self-attention can be interpreted as a gradient flow on the unit sphere, in which tokens evolve under softmax interaction potentials and tend to form clusters. While prior work has established clustering behavior for single-head attention, the multi-head setting remains less understood due to geometric interference between heads, which invalidates standard monotonicity arguments. In this work, we develop a theoretical framework for multi-head self-attention dynamics and resolve several open questions. We show that, under suitable conditions on the score matrices, a natural multi-head energy functional is non-decreasing along both flat and spherical dynamics. We identify the key obstruction to per-head monotonicity as radial shadow terms, which are projections of each head's output onto token directions, persisting even under orthogonality assumptions. We introduce a sufficient condition ensuring monotonicity and establish robustness to approximate orthogonality. In a simplified scalar-head regime with equiangular token configurations, we derive a closed-form expression for the critical inverse temperature governing clustering behavior, and show that heterogeneous heads exhibit super-additive clustering rates. In this regime, we also prove a separation in clustering time between ReLU and softmax attention in the linearized dynamics. Finally, we establish an entropy production identity and show that attention entropy increases monotonically toward equilibrium as clustering progresses. Our results provide a unified perspective on the dynamics of multi-head attention and clarify the mechanisms underlying clustering and stability in transformer models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。