arXiv:2604.01978math.PRcs.LG2026-04被引 4

研究Transformer深层注意力的随机模型,揭示其动态演化规律。

Homogenized Transformers

  • 将深度视为时间变量,构建残差流的粒子系统模型。
  • 在合适缩放下得到非平凡的均质极限,可为确定性或含共同噪声的随机过程。
  • 在高斯设定下发现表征坍塌机制,给出维度与上下文长度的量化权衡。

我们研究一种深层多头自注意力的随机模型,其中权重在各层和头之间独立重采样,类似于训练初始化状态。将深度视为时间变量,残差流在单位球面上定义了一个离散时间的相互作用粒子系统。在合适的深度、残差步长和头数联合缩放下,该动力学存在非平凡的均质极限。根据缩放方式,极限要么是确定性的,要么是带有共同噪声的随机过程;在平均场情形下,后者导出代表性标记条件分布的随机非线性福克-普朗克方程。在高斯设定下,极限漂移项消失,使均质动力学足够明确以研究表示坍塌现象。这给出了维度、上下文长度与温度之间的定量权衡,并识别出可缓解聚类的参数区域。

原文摘要 · Abstract (English)

We study a random model of deep multi-head self-attention in which the weights are resampled independently across layers and heads, as at initialization of training. Viewing depth as a time variable, the residual stream defines a discrete-time interacting particle system on the unit sphere. We prove that, under suitable joint scalings of the depth, the residual step size, and the number of heads, this dynamics admits a nontrivial homogenized limit. Depending on the scaling, the limit is either deterministic or stochastic with common noise; in the mean-field regime, the latter leads to a stochastic nonlinear Fokker--Planck equation for the conditional law of a representative token. In the Gaussian setting, the limiting drift vanishes, making the homogenized dynamics explicit enough to study representation collapse. This yields quantitative trade-offs between dimension, context length, and temperature, and identifies regimes in which clustering can be mitigated.

Transformer均质化表征坍塌动力系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。