揭示稀疏更新下动量的动态规律,解释为何全局动量在高频词上失效。
Dynamics of Stochastic Momentum with Sparse Updates in High Dimensions

- 分析稀疏输入下的动量模型,推导出二阶矩的精确解。
- 发现动量保留与学习速度的平衡决定系统稳定性和收敛性。
- 解释高频词与低频词动量冲突,为自适应动量设计提供依据。
现有动量理论假设梯度以相对恒定速率到达每个参数,但这一假设在重尾数据分布和现代架构中常被违背。本文理论分析两种可解析的稀疏更新动量模型:带有稀疏输入的最小二乘模型和存在罕见类别的逻辑回归模型。两者均具备可精确求解的二阶矩动态,并在高维极限下刻画了三种稀疏度、批量大小与动量衰减的标度指数。两类问题的相变结构由两个固有时间尺度之比决定:动量保持时间尺度(缓冲区存活的活跃更新次数)与学习时间尺度(降低平方误差所需的活跃更新次数)。当学习远慢于保持时,极限行为匹配SGD;当学习更快,系统不稳定;当两者相等,恢复经典heavy-ball动力学。不同词频下的振荡动态对应不同的最优动量值,导致全局动量在词频间产生谱冲突。
原文摘要 · Abstract (English)
Existing theory of momentum assumes that gradients arrive at every parameter at a roughly constant rate, an assumption violated in practice by heavy-tailed data distributions and modern architectures. We theoretically analyze the dynamics of two tractable models of momentum under sparse updates: a least squares model with sparse inputs and a logistic regression model with a rare class. Both admit exact closed-form second-moment dynamics whose high-dimensional limits we characterize across three scaling exponents for sparsity, batch size, and momentum decay. The phase structure on both problems is governed by the ratio of two intrinsic timescales: a momentum retention timescale (how many active updates the buffer survives) and a learning timescale (how many active updates it takes to reduce the squared error). When learning is much slower than retention, the limit matches SGD; when learning is faster, the system is unstable; where the timescales coincide, we recover classical heavy-ball dynamics. The oscillatory dynamics occur at different momentum values for different token sparsity, creating a spectral conflict for global momentum across token frequencies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。