arXiv:2511.07308cs.LG2025-11中稿 · IJCAI

将神经网络训练比作理想气体,揭示了学习率与温度的对应关系。

Can Stationary Distributions of Scale-Invariant Neural Networks Be Described by the Thermodynamics of an Ideal Gas?

  • 用热力学类比解释权重衰减下神经网络的稳定分布。
  • 理论与模拟显示熵变化与实验结果高度一致。
  • 适合研究训练动态和超参数调优的学者参考。

理解深度神经网络的训练动态仍是重大开放问题,物理启发方法提供了有前景的洞见。基于此视角,我们为带有权重衰减的随机梯度下降(SGD)在尺度不变神经网络中的稳定分布构建了一个热力学框架,该设置既反映含归一化层的实际架构,又支持理论分析。我们建立训练超参数(如学习率、权重衰减)与热力学变量(如温度、压强、体积)之间的类比。从简化的各向同性噪声模型出发,我们发现SGD动力学与理想气体行为存在紧密对应,经理论与仿真验证。扩展至神经网络训练时,框架的关键预测(如稳定熵的行为)与实验观察高度吻合。该框架为解释训练动态提供了原则性基础,或可指导未来超参数调优与学习率调度器的设计。

原文摘要 · Abstract (English)

Understanding the training dynamics of deep neural networks remains a major open problem, with physics-inspired approaches offering promising insights. Building on this perspective, we develop a thermodynamic framework to describe the stationary distributions of stochastic gradient descent (SGD) with weight decay for scale-invariant neural networks, a setting that both reflects practical architectures with normalization layers and permits theoretical analysis. We establish analogies between training hyperparameters (e.g., learning rate, weight decay) and thermodynamic variables such as temperature, pressure, and volume. Starting with a simplified isotropic noise model, we uncover a close correspondence between SGD dynamics and ideal gas behavior, validated through theory and simulation. Extending to training of neural networks, we show that key predictions of the framework, including the behavior of stationary entropy, align closely with experimental observations. This framework provides a principled foundation for interpreting training dynamics and may guide future work on hyperparameter tuning and the design of learning rate schedulers.

热力学训练动态神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。