解析两层网络在幂律数据谱下的学习规律,揭示性能随数据和模型变化的数学机制。
Analyzing Neural Scaling Laws in Two-Layer Networks with Power-Law Data Spectra
- 用统计力学方法分析单次遍历梯度下降,建模师生网络结构。
- 发现数据协方差矩阵具幂律谱时,泛化误差呈现幂律标度行为。
- 揭示了指数收敛到幂律收敛的相变现象,适合研究理论深度学习者。
神经网络的缩放定律描述了模型性能随训练数据量、模型复杂度和训练时间的变化规律,通常在多个数量级上呈现幂律行为。尽管这些规律已被广泛观测,其理论解释仍不充分。本文采用统计力学方法,在师生框架下分析单次遍历随机梯度下降,其中学生与教师均为两层神经网络。研究聚焦于泛化误差对具有幂律谱的数据协方差矩阵的响应。对于线性激活函数,推导出泛化误差的解析表达式,探讨不同学习范式并识别幂律缩放出现的条件。进一步将分析扩展至非线性激活函数在特征学习阶段的情况,研究数据协方差矩阵中幂律谱对学习动态的影响。重要的是,我们发现对称平台长度取决于数据协方差矩阵的不同特征值个数和隐藏单元数量,揭示了不同配置下平台行为。此外,结果表明当数据协方差矩阵具有幂律谱时,学习过程会从指数收敛过渡到幂律收敛。该工作为神经缩放定律提供了理论支撑,并为复杂数据结构下的学习性能优化提供洞察。
原文摘要 · Abstract (English)
Neural scaling laws describe how the performance of deep neural networks scales with key factors such as training data size, model complexity, and training time, often following power-law behaviors over multiple orders of magnitude. Despite their empirical observation, the theoretical understanding of these scaling laws remains limited. In this work, we employ techniques from statistical mechanics to analyze one-pass stochastic gradient descent within a student-teacher framework, where both the student and teacher are two-layer neural networks. Our study primarily focuses on the generalization error and its behavior in response to data covariance matrices that exhibit power-law spectra. For linear activation functions, we derive analytical expressions for the generalization error, exploring different learning regimes and identifying conditions under which power-law scaling emerges. Additionally, we extend our analysis to non-linear activation functions in the feature learning regime, investigating how power-law spectra in the data covariance matrix impact learning dynamics. Importantly, we find that the length of the symmetric plateau depends on the number of distinct eigenvalues of the data covariance matrix and the number of hidden units, demonstrating how these plateaus behave under various configurations. In addition, our results reveal a transition from exponential to power-law convergence in the specialized phase when the data covariance matrix possesses a power-law spectrum. This work contributes to the theoretical understanding of neural scaling laws and provides insights into optimizing learning performance in practical scenarios involving complex data structures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。