解析单头注意力的权重谱结构,揭示其泛化与缩放规律的本质。
Single-Head Attention in High Dimensions: A Theory of Generalization, Weights Spectra, and Scaling Laws
- 基于随机矩阵理论分析高维注意力训练过程
- 预测出低秩结构与孤立谱异常,与真实模型一致
- 发现幂律谱目标下存在分步恢复的缩放定律
训练后的注意力层表现出显著且可复现的权重谱结构,包括低秩坍缩、谱块形变及孤立的谱异常,但其成因及对泛化的影响仍不明确。本文研究在合成高维序列任务上训练的单头共享注意力层的经验风险最小化问题,使用随机矩阵理论、自旋玻璃理论和近似消息传递工具,精确刻画了高维情形下的训练与测试误差、插值与恢复阈值,以及关键与查询矩阵的谱特性。理论预测了训练后查询-键映射的完整奇异值分布,包含低秩结构与孤立谱异常,定性符合更真实Transformer中的观测结果。对于具有幂律谱的目标,证明学习过程通过逐谱恢复进行,导致幂律缩放规律的出现。
原文摘要 · Abstract (English)
Trained attention layers exhibit striking and reproducible spectral structure of the weights, including low-rank collapse, bulk deformation, and isolated spectral outliers, yet the origin of these phenomena and their implications for generalization remain poorly understood. We study empirical risk minimization in a single-head tied-attention layer trained on synthetic high-dimensional sequence tasks generated from the attention-indexed model. Using tools from random matrix theory, spin-glass theory, and approximate message passing, we obtain an exact high-dimensional characterization of training and test error, interpolation and recovery thresholds, and the spectrum of the key and query matrices. Our theory predicts the full singular-value distribution of the trained query-key map, including low-rank structure and isolated spectral outliers, in qualitative agreement with observations in more realistic transformers. Finally, for targets with power-law spectra, we show that learning proceeds through sequential spectral recovery, leading to the emergence of power-law scaling laws.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。