揭示过参数二次网络中学习能力的精确边界
The Nuclear Route: Sharp Asymptotics of ERM in Overparameterized Quadratic Networks
- 将正则化学习转化为核范数惩罚的凸矩阵感知问题
- 发现模型容量由特征映射的低秩结构决定,宽度影响可学习性
- 连接自旋玻璃、矩阵分解与凸优化,适合理论研究者
我们研究了在合成数据上训练的过参数化两层二次激活神经网络中经验风险最小化(ERM)的高维渐近行为。通过将ℓ₂正则化学习问题映射为带有核范数惩罚的凸矩阵感知任务,推导出训练误差和测试误差的精确渐近表达式。结果表明,此类网络中的容量控制源于学习到的特征映射的低秩结构。我们的分析刻画了损失函数的全局最小值,并给出了精确的泛化阈值,揭示了目标函数宽度如何决定可学习性。该研究融合并扩展了自旋玻璃方法、矩阵分解与凸优化的思想,强调了低秩矩阵感知与二次神经网络学习之间的深层联系。
原文摘要 · Abstract (English)
We study the high-dimensional asymptotics of empirical risk minimization (ERM) in over-parametrized two-layer neural networks with quadratic activations trained on synthetic data. We derive sharp asymptotics for both training and test errors by mapping the $\ell_2$-regularized learning problem to a convex matrix sensing task with nuclear norm penalization. This reveals that capacity control in such networks emerges from a low-rank structure in the learned feature maps. Our results characterize the global minima of the loss and yield precise generalization thresholds, showing how the width of the target function governs learnability. This analysis bridges and extends ideas from spin-glass methods, matrix factorization, and convex optimization and emphasizes the deep link between low-rank matrix sensing and learning in quadratic neural networks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。