arXiv:2605.10931math.APcs.LG2026-05被引 4

揭示了低温下Transformer推理时的注意力集中现象及其数学机制。

Quantifying Concentration Phenomena of Mean-Field Transformers in the Low-Temperature Regime

论文配图:Quantifying Concentration Phenomena of Mean-Field Transformers in the Low-Temperature Regime
图 1 · 摘自论文原文
  • 通过粒子系统类比,建立注意力动态的平均场方程。
  • 证明分布浓度随温度降低呈指数收敛,时间尺度为logβ。
  • 理论与实验结合,揭示值矩阵对长期行为的主导作用。

以自注意力为核心组件的Transformer已成为现代大语言模型和基础模型的关键架构。本文研究深度编码器仅在推理阶段的标记演化,其在大标记极限下由平均场连续性方程描述。借鉴多粒子相互作用系统的收敛分析思想(将标记视为粒子),我们证明:标记分布会快速集中在由键、查询、值矩阵诱导的投影映射下的初始分布的前推上,并在中等时间内保持亚稳态。具体而言,两分布间的Wasserstein距离在温度参数β⁻¹→0及推理时间t≥0下,按√(log(β+1)/β)exp(Ct)+exp(-ct)衰减。证明中,我们建立了零温方程的Lyapunov型估计,识别其t→∞时的极限,并利用Wasserstein空间中的稳定性估计与定量Laplace原理耦合两方程。结果表明,在时间尺度为logβ时,标记分布会集中于所识别的极限分布。数值实验验证了该结论,并进一步显示,对于有限β和大时间t,动态进入由值矩阵谱主导的不同终态相。

原文摘要 · Abstract (English)

Transformers with self-attention modules as their core components have become an integral architecture in modern large language and foundation models. In this paper, we study the evolution of tokens in deep encoder-only transformers at inference time which is described in the large-token limit by a mean-field continuity equation. Leveraging ideas from the convergence analysis of interacting multi-particle systems, with particles corresponding to tokens, we prove that the token distribution rapidly concentrates onto the push-forward of the initial distribution under a projection map induced by the key, query, and value matrices, and remains metastable for moderate times. Specifically, we show that the Wasserstein distance of the two distributions scales like $\sqrt{{\log(β+1)}/β}\exp(Ct)+\exp(-ct)$ in terms of the temperature parameter $β^{-1}\to 0$ and inference time $t\geq 0$. For the proof, we establish Lyapunov-type estimates for the zero-temperature equation, identify its limit as $t\to\infty$, and employ a stability estimate in Wasserstein space together with a quantitative Laplace principle to couple the two equations. Our result implies that for time scales of order $\logβ$ the token distribution concentrates at the identified limiting distribution. Numerical experiments confirm this and, beyond that, complement our theory by showing that for finite $β$ and large $t$ the dynamics enter a different terminal phase, dominated by the spectrum of the value matrix.

Transformer注意力机制平均场数学分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。