研究大学习率下线性Transformer的训练动态,发现可能陷入周期、混沌或发散而非稳定收敛。
Large-Step Training Dynamics of a Two-Factor Linear Transformer Model

- 构建可精确求解的单提示线性Transformer模型,分析其梯度下降行为。
- 在0<μ<2时,系统存在不变椭圆,分离吸引区域与混沌行为。
- 大步长会改变学习到的算法,导致非单一解的复杂动态,适用于高学习率研究者。
梯度流分析表明,简化的线性Transformer能学习上下文线性回归算法,但无法解释大学习率下梯度下降的有限步行为。受高学习率变压器不稳定性实证研究及二次回归的立方映射相图启发,我们研究了一个可精确求解的单提示线性Transformer训练问题。归一化后,动态简化为带有效步长参数μ的双因子乘积映射。在平衡截面上,该映射重现了从单调收敛到弹道收敛、周期与有界混沌非收敛、以及发散的已知标量立方相变。随后分析完整二维系统,证明当0<μ<2时,存在显式不变切比雪夫椭圆,分隔前向不变区域;该椭圆承载非平衡混沌动力,但横向排斥,而平衡标量吸引子可横向吸引。结果表明,大常数学习率会改变所学Transformer的训练吸引子,不仅加速收敛,还可能导致周期、有界混沌或发散,而非单一上下文线性回归解。我们还讨论了对基于小批量梯度下降训练方法的影响。
原文摘要 · Abstract (English)
Gradient-flow analyses show that simplified linear transformers can learn the in-context linear-regression algorithm, but they do not explain the finite-step behavior of gradient descent at large learning rates. Motivated by empirical work on high-learning-rate transformer instabilities and by the cubic-map phase diagram for quadratic regression, we study an exactly reducible one-prompt linear-transformer training problem. After normalization, the dynamics reduce to a two-factor product map with an effective step-size parameter \(μ\). On the balanced slice, this map recovers the known scalar cubic transition from monotone convergence to catapult convergence, periodic and chaotic bounded nonconvergence, and divergence. We then analyze the full two-dimensional system and show that, for \(0<μ<2\), it has an explicit invariant Chebyshev ellipse separating forward-invariant regions; this ellipse carries off-balanced chaotic dynamics but is transversely repelling, while balanced scalar attractors can be transversely attracting. These results show that large constant learning rates can change the training attractor of the learned transformer rather than merely accelerating convergence: beyond sharp stability thresholds, finite-step training may settle into cycles, bounded chaos, or divergence instead of a single in-context linear-regression solution. We also discuss the consequences for mini-batch gradient descent based training methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。