通过周期性重置外层动量,提升高维优化中通信效率与稳定性。
Outer-Momentum Restarting in High-Dimensional Two-Phase Optimization

- 引入外层动量周期性重启机制,控制局部更新的累积记忆。
- 理论证明重启可利用相位抵消,丢弃过时动量并保留有效进展。
- 适用于大规模语言模型训练,显著扩大稳定学习率范围。
通信高效的分布式优化器(如DiLoCo)通过允许工作节点在聚合前进行多次本地更新来降低同步开销。近期理论表明,外层优化器作用于由内层优化循环产生的有效谱,外层动量的选择决定了局部更新在通信周期间的积累方式。本文研究了周期性重启外层动量作为一种简单但有效的互补机制,用于控制这种外层记忆。在预测空间残差服从经验神经正则化核(NTK)的线性化平方损失模型中,我们推导出逐模式重启收缩结果,表明重启通过丢弃陈旧动量而保留内层进展,实现相位抵消。小规模实验验证了预测的收缩行为,语言模型预训练实验显示,周期性重启可拓宽外层学习率和动量值在通信周期内的稳定范围。
原文摘要 · Abstract (English)
Communication-efficient distributed optimizers such as DiLoCo reduce synchronization costs by letting workers perform many local updates before aggregating their progress with an outer momentum optimizer. Recent theory suggests that the outer optimizer acts on an effective spectrum induced by the inner optimization loop, and that the choice of outer momentum controls how progress from local updates is accumulated across communication rounds. We study periodic restarting of the outer momentum as a simple complementary mechanism for controlling this outer memory. In a linearized squared-loss model where prediction-space residuals evolve under the empirical NTK, we derive a mode-wise restart contraction showing that resets exploit phase cancellation by discarding stale momentum while preserving inner-loop progress. Toy experiments verify the predicted contraction behavior, and language-model pretraining experiments show that periodic restarts widen the stable range of outer learning rates and momentum values across communication periods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。