用热力学效应解释大模型训练中为何先升温再降速
Mpemba Effect in Large-Language Model Training Dynamics: A Minimal Analysis of the Valley-River model
- 将损失曲面比作山谷河流,快方向快速收敛,慢方向决定整体下降
- 发现高温起步反而更快降温,高平台期能加速后期衰减阶段
- 提出'强Mpemba点'理论,指导学习率调优,减少试错成本
大语言模型训练常用的学习率调度策略(暖启动、平台期、衰减)缺乏机理解释,平台高度和衰减方式多依赖经验。本文通过热力学中的Mpemba效应(加热系统比冷系统更快冷却)类比训练动态,分析一类‘山谷-河流’型损失景观:陡峭方向快速平衡,平坦方向主导全局下降。该效应解释了暖启动的必要性,并表明平台期应设为高而非低值,以加速衰减阶段的损失下降。研究发现,在特定损失景观下存在最优平台学习率——‘强Mpemba点’,此时最慢模式消失,使衰减阶段收敛更快。我们推导出其存在的解析条件,并估算维持Mpemba优势所需的衰减动力学。本最小模型分析为基于平台的学习率调度提供了原则性依据,为大模型学习率调优提供最少超参搜索指引。
原文摘要 · Abstract (English)
Learning rate (LR) schedules in large language model (LLM) training often follow empirical templates: warm-up, constant plateau/stable phase, and decay (WSD). However, the mechanistic explanation for this strategy remains underexplored, and the choice of plateau height and decay schedule is largely heuristic. In this paper, we connect training dynamics to a thermodynamic analogy via the Mpemba effect - a phenomenon in which a hotter system cools faster than a colder one when quenched into the same bath. We analyze a class of "valley-river" loss landscapes, where sharp (valley) directions equilibrate quickly, while flatter (river) directions govern global descent. The Mpemba effect provides an explanation for the necessity of the warm-up phase and motivates a high plateau - rather than a low one - for accelerating loss decrease during decay. We show that for certain loss landscapes, there exists an optimal plateau learning rate - the "strong Mpemba point" - at which the slowest mode vanishes, resulting in faster convergence during the decay phase. We derive analytical conditions for its existence and estimate decay dynamics required to preserve the Mpemba advantage. Our minimal model and analysis offer a principled justification for plateau-based schedulers and provide guidance for tuning LR in LLMs with minimal hyperparameter sweep.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。