arXiv:2606.18303cs.LGcs.AI2026-06中稿 · the 35th Internati…

揭示神经网络训练中波动与对称性降维的数学关联,解释梯度消失与突变现象。

A Link between Shock-wave Theory and Symmetry-reduced Stochastic Gradient Descent for Artificial Neural Networks

  • 通过几何与流体理论,将参数对称性降维后建模为黏性哈密顿-雅可比方程
  • 在多层感知机、Transformer等模型中,损失梯度满足类似伯格斯方程,可严格推导激波形成
  • 提出用对称性修正的观测量监控训练过程,避免原始参数失真误导

我们建立了冲击波理论与神经网络随机梯度下降对称性约化学习动力学之间的显式数学联系,基于微分几何、李群理论和流体力学。具体而言,在消除参数对称性并应用局部熵粗粒化后,有效动力学在商流形上满足黏性哈密顿-雅可比方程。若原始参数动力学可由商空间上的梯度场概括,则粗粒化损失函数的梯度服从伯格斯型方程,激波形成可严格证明。我们将该理论应用于多层感知机、卷积神经网络、Transformer及均场网络,证实其均遵循哈密顿-雅可比或伯格斯型方程。我们推测该框架可为深度学习提供实用诊断工具。在Transformer等架构中,原始参数范数常因对称冗余而扭曲,可能误导判断;而对称性修正后的商可观测量则提供了监控、预测和控制训练相变的合理基础。

原文摘要 · Abstract (English)

We develop a mathematically explicit link between shock-wave theory and the symmetry-quotiented learning dynamics of stochastic gradient descent, drawing on differential geometry, Lie group theory, and fluid mechanics. Specifically, after quotienting parameter symmetries and applying local-entropy coarse-graining, the effective dynamics satisfy a viscous Hamilton--Jacobi equation on the quotient manifold. Moreover, under the assumption that the raw parameter dynamics can be summarized by a gradient field on the quotiented space, the gradient of the coarse-grained loss function obeys a Burgers-type equation, and shock formation can be established rigorously. We apply our theory to multilayer perceptrons, convolutional neural networks, Transformers, and mean-field networks, and show that they obey the Hamilton--Jacobi or Burgers-type equations. We conjecture that this framework also yields practical diagnostics for deep learning. In architectures such as Transformers, raw parameter norms are often distorted by symmetry redundancy and may therefore be misleading, whereas symmetry-corrected quotient observables provide a principled basis for monitoring, forecasting, and controlling training-phase transitions.

深度学习对称性动力学建模激波理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。