arXiv:2607.16720cs.LG2026-07

揭示宽度相关超参数与正则化如何影响两层ReLU网络的损失曲面结构。

Effects of width-dependent model hyperparameters and $\ell_2$-regularization on the loss landscape of two-layer ReLU networks

论文配图:Effects of width-dependent model hyperparameters and $\ell_2$-regularization on the loss landscape of two-layer ReLU networks
图 1 · 摘自论文原文
  • 推导出全局最优解坍缩到零解的超参数充分条件。
  • 发现AdamW可避免参数坍缩,而SGD不能,解释其训练优势。
  • 解析证明正则化对连接性宽度无关,但降维效应随宽度增强。

理解深度神经网络仍是机器学习的核心挑战。尤其对于两层ReLU网络,在权重衰减存在时的理论性质仍不清晰。本文推导出使全局最小值坍缩至零解的超参数设置充分条件。实验显示,使用AdamW优化器可防止参数坍缩,而使用SGD则不能,这可能解释了AdamW在深度学习训练中的成功。此外,当输入维度限制为一维时,我们推导出两层ReLU网络全局最优参数集的解析解,并表明ℓ₂-正则化对连接性具有宽度不变效应,但其降维效应随网络宽度增加而增强。这些结果揭示了宽度相关超参数如何影响正则化损失曲面的几何结构。

原文摘要 · Abstract (English)

Understanding deep neural networks remains a central challenge in machine learning. In particular, the theoretical properties of even two-layer ReLU networks, especially in the presence of weight decay, remain poorly understood. To this end, we derive a sufficient condition on the hyperparameter settings under which the global minima collapse to the zero solution. Interestingly, our experiments reveal that using AdamW as an optimizer prevents the collapse of the learned parameters, whereas using SGD does not, which may help explain the success of AdamW in deep learning training. In addition, when restricting the input dimension to one, we derive an analytical solution for the globally optimal parameter sets of two-layer ReLU networks and show that $\ell_2$-regularization has a width-invariant effect on connectivity, but its dimensionality-reducing effect becomes stronger as the network width increases. These results provide insight into how width-dependent hyperparameters influence the geometry of regularized loss landscapes.

深度学习损失曲面正则化两层网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。