揭示注意力机制中复制子电路的突现原理,解释为何训练数据量达到临界点时会突然出现。
Phase Transitions in Attention: A Bayesian Theory of Copy Head Emergence

- 基于贝叶斯推导注意力矩阵后验,将高维问题简化为低维序参量空间。
- 发现训练数据量达到临界值时,复制子电路出现一级相变,非渐进式演化。
- 对比线性注意力,说明软最大注意力具有突变特性,适合理解大模型训练现象。
注意力是变压器模型上下文学习的核心机制,其注意力模式在训练中被观察到会突然涌现。本文提出一种注意力特征学习的贝叶斯理论,并聚焦于第一层归纳头中的复制子电路如何通过单层softmax注意力网络在复制任务上学习。我们推导出注意力矩阵的闭式后验,并将其简化为低维序参量空间。该简化揭示了训练数据量存在相变点,通过贝叶斯采样和标准Adam训练均验证了这一现象。与线性注意力对比发现,软最大注意力表现出一阶相变,而线性注意力初始为二阶相变,随后平滑演变为结构化注意力模式(交叉过渡)。本工作首次从第一性原理提供了复制子电路突现的理论解释,与大型语言模型训练中观测到的现象相似。
原文摘要 · Abstract (English)
Attention is the key mechanism underlying in-context learning in transformers, and attention patterns have been observed empirically to emerge abruptly during training. We present a Bayesian theory of feature learning in attention; we then focus on how the copy subcircuit in the first layer of an induction head is learned by analyzing a single-layer softmax attention network trained on a copy task. We derive a closed-form posterior over the attention matrix and reduce it to a low-dimensional order parameter space. This reduction reveals a phase transition in the amount of training data, which we verify using both Bayesian sampling and standard training with Adam. We contrast our results with linear attention and find that softmax attention exhibits a \emph{first-order phase transition} while in linear attention an initial \emph{second-order phase transition} is followed by a smooth, continuous evolution toward the structured attention pattern (\emph{crossover}). Our work provides a first-principles theoretical account of the abrupt emergence of the copy subcircuit, reminiscent of the one observed in training large language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。