arXiv:2508.12837cs.LG2025-08ICML被引 17

揭示Transformer学习n元语法时,子n元语法接近稳定点的理论机制。

Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points

  • 通过构造简化模型中的参数配置,证明子n元语法是损失函数的近似驻点。
  • 在序列无限长和参数范数大的极限下,这些点的梯度趋近于零。
  • 解释了训练中阶段性进展与相变现象,适合研究训练动态的读者。

受训练过程中长期平台期和分阶段进展现象的启发,我们研究了在上下文内进行下一个词预测任务的Transformer模型的损失曲面。特别地,针对在交叉熵损失下学习上下文n元语法语言模型的问题,建立了参数配置成为驻点的充分条件。我们构造了一类简化Transformer模型的参数配置,用于表示k元语法估计器(k ≤ n),并证明在序列长度和参数范数趋于无穷时,这些解处的总体损失梯度消失。这揭示了损失曲面的关键性质:子n元语法是总体交叉熵损失的近似驻点,为广泛观察到的分阶段学习动态和涌现相变现象提供了理论解释。数值实验进一步展示了n元语法的学习动态,表现为在近似驻点间的离散跃迁。

原文摘要 · Abstract (English)

Motivated by empirical observations of prolonged plateaus and stage-wise progression during training, we investigate the loss landscape of transformer models trained on in-context next-token prediction tasks. In particular, we focus on learning in-context $n$-gram language models under cross-entropy loss, and establish a sufficient condition for parameter configurations to be stationary points. We then construct a set of parameter configurations for a simplified transformer model that represent $k$-gram estimators (for $k \leq n$), and show that the gradient of the population loss at these solutions vanishes in the limit of infinite sequence length and parameter norm. This reveals a key property of the loss landscape: {sub-$n$-grams are near-stationary points of the population cross-entropy loss}, offering theoretical insight into widely observed phenomena such as stage-wise learning dynamics and emergent phase transitions. These insights are further supported by numerical experiments that illustrate the learning dynamics of $n$-grams, characterized by discrete transitions between near-stationary solutions.

Transformer损失曲面n元语法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。