arXiv:2603.03993cs.LGcond-mat.dis-nn2026-03被引 1

解析Transformer注意力头的分工机制,揭示其训练中逐步特化的规律。

Specialization of softmax attention heads: insights from the high-dimensional single-location model

  • 构建高维单位置模型,解释注意力头如何分阶段对齐潜在信号方向
  • 发现训练初期无分工,后期各头依次聚焦不同语义方向
  • 提出贝叶斯Softmax注意力,理论性能最优,适合研究注意力机制

多头注意力使Transformer模型能同时表示多种注意力模式。实证发现,注意力头在训练过程中分阶段出现分工现象,而许多头仍冗余且学习相似表示。本文提出一个基于多指标与单位置回归框架的理论模型,分析了在随机梯度下降(SGD)下多头Softmax注意力的训练动态,揭示初始未分工阶段后,不同注意力头会依次对齐潜在信号方向,进入多阶段分工过程。第二部分研究注意力激活函数对性能的影响,提出贝叶斯Softmax注意力,在该设定下实现最优预测性能。

原文摘要 · Abstract (English)

Multi-head attention enables transformer models to represent multiple attention patterns simultaneously. Empirically, head specialization emerges in distinct stages during training, while many heads remain redundant and learn similar representations. We propose a theoretical model capturing this phenomenon, based on the multi-index and single-location regression frameworks. In the first part, we analyze the training dynamics of multi-head softmax attention under SGD, revealing an initial unspecialized phase followed by a multi-stage specialization phase in which different heads sequentially align with latent signal directions. In the second part, we study the impact of attention activation functions on performance. We introduce the Bayes-softmax attention, which achieves optimal prediction performance in this setting.

注意力机制Transformer理论分析深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。