用重整化群理论揭示注意力机制在不同数据下是相关还是无关
Relevant and Irrelevant: A Renormalization Group Analysis of Transformer Attention
- 将注意力视为对MLP固定点的微扰,用重整化群分析其相关性
- 长相关数据中注意力显著提升表示维度,短相关则无影响
- 首层注意力头主导变化,适合研究数据结构与模型行为关系者
基于威尔逊重整化群理论,我们将Transformer的注意力机制视为对训练后MLP残差堆栈固定点的微扰,并探究其是否为相关、边缘或无关算子。推导出固定点偏移公式,得出四项可验证预测:固定点几何、有效秩谱、层特异性及微扰衰减谱。在具有可控相关长度的合成马尔可夫链序列上测试发现:(1) 长相关序列中,注意力强相关——填补了MLP无法弥合的残差损失缺口,引发表征空间相变,有效秩在第1层跃过输入维度并稳定于高维平台;(2) 短相关序列中,注意力无关——模型收敛至与MLP相同的损失和固定点几何,但更快压缩微扰;(3) 相变主要由第一层第一个注意力头(L0H0)主导,其表征偏移量超过后续任意头的4倍,符合相关算子作用于位置信息整合前的预测;(4) 微扰衰减实验显示谱选择性反转:长相关下,Transformer选择保留缓慢马尔可夫模式(衰减长度动态范围5.4倍于MLP,后者为1.3倍);短相关下则抑制所有模式快于MLP,无谱选择性。结果表明,注意力的相关性并非架构固有属性,而是数据生成过程谱结构的函数,且一阶重整化群框架能有效预测此差异。
原文摘要 · Abstract (English)
Using the language of Wilsonian renormalization group theory (RG), we treat the Transformer's attention mechanism as a perturbation of the trained MLP residual-stack fixed point and ask whether it constitutes a relevant, marginal, or irrelevant operator. We derive a fixed-point shift formula and obtain four testable predictions for the fixed-point geometry, effective rank profile, layer specificity, and perturbation decay spectrum. Testing these on synthetic Markov chain sequences with controlled correlation length, we find: (1) For large chains(long correlation), attention is strongly relevant: it closes a residual loss gap the MLP cannot bridge and drives a phase transition in representation space, with effective rank jumping above input dimensionality at layer 1 and stabilizing at a high-dimensional plateau. (2) For short chains(short correlation), attention is irrelevant: the Transformer converges to the same loss and fixed-point geometry as the MLP, though it contracts perturbations faster. (3) The transition is dominated by the first-layer head (L0H0), which accounts for more than 4 times the representational shift of any subsequent head, consistent with the prediction that the relevant operator acts before the MLP begins integrating out positional variation. (4) Perturbation decay experiments reveal a regime reversal: in the long correlation regime the Transformer selectively preserves slow Markov modes (5.4 times the dynamic range in decay length vs. 1.3 times for the MLP); in the short correlation regime it suppresses all modes faster than the MLP, with no spectral selectivity. Together, these results show that the relevance of attention is not a property of the architecture but of the spectral structure of the data-generating process, and that a first-order RG perturbation framework provides a predictive account of that difference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。