arXiv:2604.25150cs.LGcs.AI2026-04

揭示过参数化如何通过对称性改善神经网络优化

The Role of Symmetry in Optimizing Overparameterized Networks

论文配图:The Role of Symmetry in Optimizing Overparameterized Networks
图 1 · 摘自论文原文
  • 分析权重空间对称性,发现过参数化引入新对称
  • 对称性使最优解更易到达,降低损失曲面条件数
  • 适合研究深度学习优化机制的读者

过参数化是深度学习成功的核心,但其如何促进优化仍不明确。本文分析神经网络中的权重空间对称性,表明过参数化引入额外对称性,从两方面提升优化:第一,这些对称性在海森矩阵上起到类似对角预处理的作用,使得每个函数等价解类中存在条件更优的极小值;第二,过参数化提高了全局最小值在典型初始化附近的概率密度,使其更易被找到。实验验证了更宽的网络具有更低的主特征值、更小的条件数和更快的收敛速度,与理论分析一致。本研究为损失曲面几何与简单性偏好之间提供了潜在联系,并构建了一个统一框架,将过参数化与宽度增长理解为损失曲面的几何变换。

原文摘要 · Abstract (English)

Overparameterization is central to the success of deep learning, yet the mechanisms by which it improves optimization remain incompletely understood. We analyze weight-space symmetries in neural networks and show that overparameterization introduces additional symmetries that benefit optimization in two distinct ways. First, we prove that these symmetries act as a form of diagonal preconditioning on the Hessian, enabling the existence of better-conditioned minima within each equivalence class of functionally identical solutions. Second, we show that overparameterization increases the probability mass of global minima near typical initializations, making these favourable solutions more reachable. These results offer a potential link between loss landscape geometry and simplicity bias. Empirically, we observe wider networks have lower top eigenvalues, smaller condition numbers and faster convergence, matching our analysis. Our analysis provides a unified framework for understanding overparameterization and width growth as a geometric transformation of the loss landscape.

神经网络优化对称性过参数化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。