揭示高维两层ReLU网络局部极小值的精确结构及其与梯度下降的关系。
Sharp description of local minima in the loss landscape of high-dimensional two-layer ReLU neural networks
- 用统计量精确描述局部极小值,实现低维表征。
- 局部极小对应梯度下降的吸引不动点,且过参数化提升全局最优解可达性。
- 揭示常见简化假设的局限性,适用于理解神经网络优化本质的研究者。
我们研究在高斯协变量的可实现教师-学生设定下,形如∑_{k=1}^K ReLU(w_k^⊤x) 的两层ReLU网络的总体损失景观。我们证明局部极小值可由总结统计量精确地低维表示,从而实现对景观的清晰可解释刻画。进一步建立了与单次遍历随机梯度下降(SGD)的直接联系:局部极小对应于总结统计量空间中吸引的不动点。这一视角揭示了极小值的分层组织结构,并表明过参数化会改变其稳定性和梯度动力学下的可达性。在过参数化情形下,全局极小值变得越来越易被访问,主导动态过程并减少收敛至虚假解的可能性。总体而言,我们的结果揭示了常见简化假设的内在局限性,即使在最简化的神经网络模型中也可能遗漏关键损失景观特征。
原文摘要 · Abstract (English)
We study the population loss landscape of two-layer ReLU networks of the form $\sum_{k=1}^K \mathrm{ReLU}(w_k^\top x)$ in a realisable teacher-student setting with Gaussian covariates. We show that local minima admit an exact low-dimensional representation in terms of summary statistics, yielding a sharp and interpretable characterisation of the landscape. We further establish a direct link with one-pass SGD: local minima correspond to attractive fixed points of the dynamics in summary statistics space. This perspective reveals a hierarchical organisation of minima into discrete families and shows how overparameterisation changes their stability and reachability under gradient-based dynamics. In this overparameterised regime, global minima become increasingly accessible, attracting the dynamics and reducing convergence to spurious solutions. Overall, our results reveal intrinsic limitations of common simplifying assumptions, which may miss essential features of the loss landscape even in minimal neural network models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。