揭示注意力模型高维训练中的动态机制与隐式偏置
High-Dimensional Learning Dynamics of Attention-Indexed Models

- 构建注意力索引模型框架,分析高维下损失景观与梯度演化
- 发现固定参数需约 $\Theta(d^2\log d)$ 样本才能弱恢复
- 揭示绑定与解绑注意力的对称性破缺差异,适合理解大模型训练
注意力机制是现代基础模型的核心,但其训练动态在注意力矩阵高秩时仍不清晰。本文研究注意力索引模型框架,涵盖多层多头结构。在高维极限下,群体损失景观由有限个迹序参数刻画;而在线随机梯度下降(SGD)受无限阶矩阵矩支配,可被有限截断系统指数逼近。该框架表明,注意力参数化本身即具架构隐式偏置:直接优化注意力矩阵 $S\in\mathbb{R}^{d\times d}$ 易陷入无信息状态;绑定注意力($S=WW^\top$)自动引发对称性破缺,可在 $\Theta(d^2\log d)$ 样本下实现弱恢复。对解绑注意力($S=UV^\top$),发现快慢机制:预激活均值快速演化,重叠缓慢变化;当快速动力学选择的状态打破初始对称性时,可在 $\Theta(d^2\log d)$ 样本量下实现弱恢复。
原文摘要 · Abstract (English)
Attention mechanisms are central to modern foundation models, yet their training dynamics remain poorly understood, especially when the attention matrices have extensive rank. In this work, we study attention-indexed models, a broad framework that can represent multi-layer and multi-head attention architectures. First, we show that, in a suitable high-dimensional limit, the population-loss landscape is characterized by a finite set of trace order parameters. In contrast, online stochastic gradient descent (SGD) is governed by an infinite hierarchy of matrix moments, which we show can be exponentially well-approximated by a finite truncated system. Second, this framework reveals that attention parameterization itself can act as an architectural implicit bias. Direct optimization of an attention matrix $S\in\mathbb{R}^{d\times d}$ can remain trapped in an uninformative state. Tied attention ($S=WW^\top$) induces an automatic symmetry-breaking mechanism and yields weak recovery in $\Theta(d^2\log d)$ samples. For untied attention, $S=UV^\top$, we uncover a fast-slow mechanism: the pre-activation mean first evolves on a fast timescale, while the overlaps evolve on a slower one. Weak recovery on the $\Theta(d^2\log d)$ scale occurs when the state selected by the fast dynamics breaks the initial symmetry.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。