揭示浅层网络在特征学习中的缩放规律与权重谱的关系
Scaling Laws and Spectra of Shallow Neural Networks in the Feature Learning Regime
- 通过矩阵压缩感知理论分析二次与对角神经网络的缩放行为
- 发现过拟合与泛化间的相变点及风险下降的平台期现象
- 首次从原理上解释权重谱幂律尾部与泛化性能的关联
深度学习中的缩放定律推动了近年来诸多进展,但其理论理解仍主要局限于线性模型。本文系统分析了在特征学习范式下二次与对角神经网络的缩放规律。借助矩阵压缩感知与LASSO的联系,推导出过剩风险缩放指数随样本复杂度和权重衰减变化的详细相图。该分析揭示了不同缩放阶段之间的转变以及平台行为,与经验研究中广泛报告的现象一致。此外,我们建立了这些阶段与训练后网络权重谱性质的精确关联,并进行了详尽刻画。由此,为近期经验观察——即权重谱中幂律尾部的出现与网络泛化性能相关——提供了首原则解释。
原文摘要 · Abstract (English)
Neural scaling laws underlie many of the recent advances in deep learning, yet their theoretical understanding remains largely confined to linear models. In this work, we present a systematic analysis of scaling laws for quadratic and diagonal neural networks in the feature learning regime. Leveraging connections with matrix compressed sensing and LASSO, we derive a detailed phase diagram for the scaling exponents of the excess risk as a function of sample complexity and weight decay. This analysis uncovers crossovers between distinct scaling regimes and plateau behaviors, mirroring phenomena widely reported in the empirical neural scaling literature. Furthermore, we establish a precise link between these regimes and the spectral properties of the trained network weights, which we characterize in detail. As a consequence, we provide a theoretical validation of recent empirical observations connecting the emergence of power-law tails in the weight spectrum with network generalization performance, yielding an interpretation from first principles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。