揭示自注意力矩阵的谱特性,证明其渐近等价于高斯模型。
Gaussian Equivalence for Self-Attention: Asymptotic Spectral Analysis of Attention Matrix
- 基于随机矩阵理论,分析注意力矩阵的奇异值分布。
- 发现平方奇异值偏离马尔琴科-帕斯图律,与以往认知不同。
- 适用于研究注意力机制理论性质的算法与理论研究人员。
自注意力层已成为现代深度神经网络的核心组件,但其理论理解仍不充分,尤其缺乏随机矩阵理论视角下的分析。本文对注意力矩阵的奇异值谱进行了严格分析,首次建立了注意力的高斯等价性结果。在逆温度保持常数的自然设定下,我们证明注意力矩阵的奇异值分布可由一个可解析的线性模型渐近刻画。进一步发现,平方奇异值分布偏离了此前文献中广为接受的马尔琴科-帕斯图律。证明依赖两个关键要素:对归一化项波动的精确控制,以及利用指数函数有利泰勒展开的精细线性化。该分析还识别出线性化的阈值,并阐明尽管注意力并非逐元素操作,但在该设定下仍能实现严格的高斯等价性。
原文摘要 · Abstract (English)
Self-attention layers have become fundamental building blocks of modern deep neural networks, yet their theoretical understanding remains limited, particularly from the perspective of random matrix theory. In this work, we provide a rigorous analysis of the singular value spectrum of the attention matrix and establish the first Gaussian equivalence result for attention. In a natural regime where the inverse temperature remains of constant order, we show that the singular value distribution of the attention matrix is asymptotically characterized by a tractable linear model. We further demonstrate that the distribution of squared singular values deviates from the Marchenko-Pastur law, which has been believed in previous work. Our proof relies on two key ingredients: precise control of fluctuations in the normalization term and a refined linearization that leverages favorable Taylor expansions of the exponential. This analysis also identifies a threshold for linearization and elucidates why attention, despite not being an entrywise operation, admits a rigorous Gaussian equivalence in this regime.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。