arXiv:2603.06248cs.LGmath.OC2026-03

梯度流让softmax输出趋向低熵解,解释了Transformer的训练现象。

Gradient Flow Polarizes Softmax Outputs towards Low-Entropy Solutions

  • 分析值-softmax模型的梯度流,揭示其天然倾向低熵解。
  • 在逻辑损失和平方损失下均出现输出极化现象,熵显著降低。
  • 为注意力汇聚、激活值爆炸等现象提供理论解释,适合研究自注意机制者阅读。

理解基于softmax模型的复杂非凸训练动态,对于解释Transformer的实证成功至关重要。本文分析了值-softmax模型 ${L}(\mathbf{V} σ(\mathbf{a}))$ 的梯度流动力学,其中 $\mathbf{V}$ 和 $\mathbf{a}$ 分别为可学习的值矩阵与注意力向量。由于矩阵乘softmax向量的参数化构成自注意力的核心组件,本分析直接揭示了Transformer的训练动态。我们发现,该结构上的梯度流天然驱动优化朝向低熵输出解。我们在多种目标函数(包括逻辑损失和平方损失)中验证了这一极化效应的普适性。此外,我们讨论了这些理论结果的实际意义,为注意力汇聚和大量激活等经验现象提供了形式化机制。

原文摘要 · Abstract (English)

Understanding the intricate non-convex training dynamics of softmax-based models is crucial for explaining the empirical success of transformers. In this article, we analyze the gradient flow dynamics of the value-softmax model, defined as ${L}(\mathbf{V} σ(\mathbf{a}))$, where $\mathbf{V}$ and $\mathbf{a}$ are a learnable value matrix and attention vector, respectively. As the matrix times softmax vector parameterization constitutes the core building block of self-attention, our analysis provides direct insight into transformer's training dynamics. We reveal that gradient flow on this structure inherently drives the optimization toward solutions characterized by low-entropy outputs. We demonstrate the universality of this polarizing effect across various objectives, including logistic and square loss. Furthermore, we discuss the practical implications of these theoretical results, offering a formal mechanism for empirical phenomena such as attention sinks and massive activations.

Transformer梯度流注意力机制低熵

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。