特征学习能显著提升复杂任务的模型扩展效率。
How Feature Learning Can Improve Neural Scaling Laws
- 区分任务难易度,发现特征学习仅对困难任务有效。
- 在困难任务上,训练时间与计算量的扩展指数几乎翻倍。
- 适合研究模型可扩展性或复杂任务建模的学者参考。
我们构建了一个超越核极限的神经网络扩展规律可解模型。理论分析表明,性能随模型规模、训练时间和可用数据总量的变化规律取决于任务难度,可分为硬、易和超易三类。对于位于初始无限宽神经正切核(NTK)再生核希尔伯特空间(RKHS)内的易和超易任务,特征学习与核模式的扩展指数一致。而对于不在初始NTK-RKHS中的困难任务,我们通过理论与实证证明,特征学习能显著改善训练时间与计算资源的扩展效率,使扩展指数近乎翻倍。这导致了特征学习阶段应采用不同的参数与训练时间扩展策略。我们在非线性MLP拟合圆上幂律傅里叶谱函数及CNN完成视觉任务的实验中验证了该结论:特征学习仅提升困难任务的扩展性能,对易任务无影响。
原文摘要 · Abstract (English)
We develop a solvable model of neural scaling laws beyond the kernel limit. Theoretical analysis of this model shows how performance scales with model size, training time, and the total amount of available data. We identify three scaling regimes corresponding to varying task difficulties: hard, easy, and super easy tasks. For easy and super-easy target functions, which lie in the reproducing kernel Hilbert space (RKHS) defined by the initial infinite-width Neural Tangent Kernel (NTK), the scaling exponents remain unchanged between feature learning and kernel regime models. For hard tasks, defined as those outside the RKHS of the initial NTK, we demonstrate both analytically and empirically that feature learning can improve scaling with training time and compute, nearly doubling the exponent for hard tasks. This leads to a different compute optimal strategy to scale parameters and training time in the feature learning regime. We support our finding that feature learning improves the scaling law for hard tasks but not for easy and super-easy tasks with experiments of nonlinear MLPs fitting functions with power-law Fourier spectra on the circle and CNNs learning vision tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。