arXiv:2503.06001stat.MLcs.LG2025-03中稿 · AISTATS 2025被引 4

研究神经网络权重排列不变性如何影响模型连接的平滑性。

Analyzing the Role of Permutation Invariance in Linear Mode Connectivity

  • 在两层ReLU网络中分析排列不变性下的线性模式连通性。
  • 网络宽度增大时,连接障碍呈双下降趋势,最终以O(m⁻¹/²)速率趋近于零。
  • 结果揭示排列不变性可显著降低连接障碍,适合关注模型融合的研究者。

Entezari等人(2021)实证发现,考虑神经网络的排列不变性后,两个SGD解之间沿线性插值路径几乎不存在损失屏障,即排列不变下的线性模式连通性(LMC)。本文在教师-学生设定下,对两层ReLU网络进行细粒度分析,证明当学生网络宽度$m$增加时,模排列的LMC损失屏障呈现双下降行为。特别地,当$m$足够大时,屏障以$O(m^{-1/2})$速率降至零,且该速率不受维度灾难影响,表明排列不变性能有效降低连接障碍。此外,我们观察到学习率提高时梯度下降/随机梯度下降解的稀疏性出现突变,并研究其对模排列的LMC损失屏障的影响。在合成数据和MNIST上的实验验证了理论预测,且在更复杂架构中也表现出类似趋势。

原文摘要 · Abstract (English)

It was empirically observed in Entezari et al. (2021) that when accounting for the permutation invariance of neural networks, there is likely no loss barrier along the linear interpolation between two SGD solutions -- a phenomenon known as linear mode connectivity (LMC) modulo permutation. This phenomenon has sparked significant attention due to both its theoretical interest and practical relevance in applications such as model merging. In this paper, we provide a fine-grained analysis of this phenomenon for two-layer ReLU networks under a teacher-student setup. We show that as the student network width $m$ increases, the LMC loss barrier modulo permutation exhibits a double descent behavior. Particularly, when $m$ is sufficiently large, the barrier decreases to zero at a rate $O(m^{-1/2})$. Notably, this rate does not suffer from the curse of dimensionality and demonstrates how substantial permutation can reduce the LMC loss barrier. Moreover, we observe a sharp transition in the sparsity of GD/SGD solutions when increasing the learning rate and investigate how this sparsity preference affects the LMC loss barrier modulo permutation. Experiments on both synthetic and MNIST datasets corroborate our theoretical predictions and reveal a similar trend for more complex network architectures.

模式连通性神经网络排列不变性双下降

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。