破解注意力层无限宽极限下的非高斯分布规律。
Infinite-Width Limit of a Single Attention Layer: Analysis via Tensor Programs
- 基于张量程序框架,推导单层注意力在真实维度下的极限分布。
- 发现极限分布非高斯,其结构依赖于随机相似度得分的条件分布。
- 理论适用于有限头数和标准缩放,为深层Transformer统一理论奠基。
在现代神经网络的理论分析中,无限宽极限常用于解释神经元预激活的高斯近似(如神经网络高斯过程或张量程序)。然而,这些基于高斯的渐近理论至今无法捕捉注意力层的行为,除非在特殊情形下(如无穷多头或定制缩放方案)。本文利用张量程序框架,严格确定了在真实架构维度与标准 $1/ ext{sqrt}{n}$ 缩放下,单个注意力层变量的无限宽极限分布。我们无需依赖无穷头近似或定制缩放,推导出该极限分布的精确形式,证明其根本上偏离高斯性。该分布具有层级结构,表现为在随机相似度得分条件下的高斯性。数值实验验证了理论预测的有效性,表明理论在有限宽度下依然准确,能精确描述有限头注意力。除刻画独立注意力层外,本研究为构建深度Transformer架构在无限宽极限下的统一理论奠定基础。
原文摘要 · Abstract (English)
In modern theoretical analyses of neural networks, the infinite-width limit is often invoked to justify Gaussian approximations of neuron preactivations (e.g., via neural network Gaussian processes or Tensor Programs). However, these Gaussian-based asymptotic theories have so far been unable to capture the behavior of attention layers, except under special regimes such as infinitely many heads or tailored scaling schemes. In this paper, leveraging the Tensor Programs framework, we rigorously identify the infinite-width limit distribution of variables within a single attention layer under realistic architectural dimensionality and standard $1/\sqrt{n}$-scaling with $n$ dimensionality. We derive the exact form of this limit law without resorting to infinite-head approximations or tailored scalings, demonstrating that it departs fundamentally from Gaussianity. This limiting distribution exhibits non-Gaussianity from a hierarchical structure, being Gaussian conditional on the random similarity scores. Numerical experiments validate our theoretical predictions, confirming the effectiveness of our theory at finite width and accurate description of finite-head attentions. Beyond characterizing a standalone attention layer, our findings lay the groundwork for developing a unified theory of deep Transformer architectures in the infinite-width regime.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。