揭示现代神经网络的梯度流守恒定律,解释过参数化模型的成功机制。
Conservation Laws for Modern Neural Architectures
- 构建统一框架分析GELU、SiLU等激活函数的守恒律
- 实验验证了多种架构下的理论预测不变量
- 适合研究深度学习动力学与模型泛化的学者
理解梯度下降的动力学是解释过参数化模型成功的关键,其中隐式偏差通过梯度流中的守恒定律体现。尽管线性网络和ReLU网络的这类守恒律已有较好理解,但现代架构仍缺乏系统研究。本文提出统一框架,刻画包含GELU、SiLU、SwiGLU激活函数的前馈网络、使用正弦或旋转位置编码的多头注意力,以及不同门控设计下的专家混合(Mixture-of-Experts)架构的守恒律。理论结果得到实验验证,确认了预测的不变量在实际训练中成立。
原文摘要 · Abstract (English)
Understanding gradient descent dynamics is key to explaining the success of over-parameterized models, where implicit bias manifests through conservation laws in gradient flow. While such laws are well understood for linear and ReLU networks, they remain largely unexplored for modern architectures. This work develops a unified framework to characterize conservation laws for contemporary models, including feedforward networks with GELU, SiLU, and SwiGLU activations, multihead attention with sinusoidal and rotary positional encodings, and Mixture-of-Experts architectures under diverse gating designs. Our theoretical findings are supported by experiments that validate the predicted invariants.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。