arXiv:2505.10559cs.LGcs.AI2025-05被引 12

用热力学原理揭示大模型训练规律,指导学习率设计。

Neural Thermodynamic Laws for Large Language Model Training

  • 将损失曲面类比为山谷,引入温度、熵等热力学概念。
  • 推导出热容量与学习率调度的内在关系,提升训练效率。
  • 适合对训练机制本质感兴趣的算法研究者和优化工程师。

超越神经网络缩放定律,人们对大语言模型(LLM)训练背后的规律知之甚少。本文提出神经热力学定律(NTL)——一个新框架,为LLM训练动态提供全新视角。理论上,在河谷型损失景观假设下,温度、熵、热容、热传导等关键热力学量及热力学三定律、能量均分定理等经典原理自然涌现。实践上,这一科学视角为设计学习率调度提供了直观指导。

原文摘要 · Abstract (English)

Beyond neural scaling laws, little is known about the laws underlying large language models (LLMs). We introduce Neural Thermodynamic Laws (NTL) -- a new framework that offers fresh insights into LLM training dynamics. On the theoretical side, we demonstrate that key thermodynamic quantities (e.g., temperature, entropy, heat capacity, thermal conduction) and classical thermodynamic principles (e.g., the three laws of thermodynamics and the equipartition theorem) naturally emerge under river-valley loss landscape assumptions. On the practical side, this scientific perspective yields intuitive guidelines for designing learning rate schedules.

大模型训练热力学学习率调度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。