解析注意力模型训练过程中的梯度下降动态,揭示序列结构如何影响学习效率。
Asymptotics of SGD in Sequence-Single Index Models and Single-Layer Attention Networks
- 构建序列单指标模型,用统计量刻画语义与位置对齐
- 发现训练分逃逸初始化和目标对齐两个阶段,序列长度影响收敛速度
- 为注意力机制的可解释性提供理论支持,适合研究模型动力学者
我们研究了一类称为序列单指标(SSI)模型的序列模型中随机梯度下降(SGD)的动力学行为,其中目标依赖于输入空间中的单一方向作用于一系列标记。该设定将经典单指标模型推广到序列领域,涵盖简化的一层注意力架构。我们推导出以一对充分统计量表示的总体损失闭式表达,该统计量捕捉语义与位置对齐,并刻画了这些坐标的高维SGD动态。分析揭示了两个不同的训练阶段:从无信息初始化中逃逸及与目标子空间对齐;并展示了序列长度和位置编码如何影响收敛速度与学习轨迹。这些结果为理解基于注意力模型中数据序列结构的有益性提供了严格且可解释的理论基础。
原文摘要 · Abstract (English)
We study the dynamics of stochastic gradient descent (SGD) for a class of sequence models termed Sequence Single-Index (SSI) models, where the target depends on a single direction in input space applied to a sequence of tokens. This setting generalizes classical single-index models to the sequential domain, encompassing simplified one-layer attention architectures. We derive a closed-form expression for the population loss in terms of a pair of sufficient statistics capturing semantic and positional alignment, and characterize the induced high-dimensional SGD dynamics for these coordinates. Our analysis reveals two distinct training phases: escape from uninformative initialization and alignment with the target subspace, and demonstrates how the sequence length and positional encoding influence convergence speed and learning trajectories. These results provide a rigorous and interpretable foundation for understanding how sequential structure in data can be beneficial for learning with attention-based models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。