揭示SGD优化中子网络合并的相变机制,发现方差突增现象。
Percolation Dynamics in Optimization : Variance Cascades and Discrete Scale Invariance

- 将SGD视为渗流过程,子网络以离散块形式同步合并。
- 观测到宏观序参量的方差突增,类比物理相变。
- 该机制在Adam/AdamW中同样存在,适用于重尾噪声场景。
我们研究了随机梯度下降(SGD)的动力学行为,已知其会引导深度神经网络进入对应于更简单子网络的不变集。这一引导过程随时间如何展开仍不明确。本文通过将随机梯度流(SGF)建模为渗流过程,揭示了架构对称性迫使子网络以离散同时块的形式合并,而非逐个进行。这些结构转变在宏观序参量中表现为方差尖峰,类似于物理相变。进一步表明,这种捕获机制及其关联的标度级联在显式重尾噪声模型下也适用于Adam和AdamW。
原文摘要 · Abstract (English)
We study the dynamics of Stochastic Gradient Descent (SGD), which is known to steer deep neural networks toward invariant sets that correspond to simpler subnetworks. How this steering unfolds over time remains poorly understood. We answer this by modeling the stochastic gradient flow (SGF) as a percolation process, in which architectural symmetries force subnetworks to merge in discrete simultaneous blocks rather than one at a time. These structural transitions register as variance spikes in a macroscopic order parameter, echoing physical phase transitions. We further show this trapping mechanism and its associated scaling cascade extend to Adam and AdamW under an explicit heavy-tailed noise model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。