统一解释神经网络缩放规律,揭示模型、数据、算力的协同作用机制
Effective Frontiers: A Unification of Neural Scaling Laws
- 将学习任务抽象为长尾分布模式的渐进覆盖过程
- 提出有效边界概念,量化可学习知识与未学尾部的分界
- 揭示不同缩放定律本质是资源约束下的最优平衡解
神经网络缩放定律描述了测试损失随模型容量(N)、数据量(D)和计算量(C)提升的幂律改善。然而现有理论解释多依赖特定架构或复杂核方法,缺乏直观普适性。本文提出统一框架,将一般学习任务视为从长尾(齐普夫型)分布中渐进覆盖模式的过程。引入有效边界(k⋆),即模式秩空间中的阈值,用于区分已学习知识与未学习尾部。我们证明可约损失在渐近意义上由资源依赖的边界截断所对应的尾部概率质量决定。基于该框架,我们推导出关于N、D、C的精确缩放定律,分别归因于容量、覆盖和优化瓶颈。此外,通过最大瓶颈原理统一三者机制,表明卡普兰与钦奇拉缩放定律并非矛盾,而是不同活跃瓶颈下同一受限优化问题的均衡解。
原文摘要 · Abstract (English)
Neural scaling laws govern the prediction power-law improvement of test loss with respect to model capacity ($N$), datasize ($D$), and compute ($C$). However, existing theoretical explanations often rely on specific architectures or complex kernel methods, lacking intuitive universality. In this paper, we propose a unified framework that abstracts general learning tasks as the progressive coverage of patterns from a long-tail (Zipfian) distribution. We introduce the Effective Frontier ($k_\star$), a threshold in the pattern rank space that separates learned knowledge from the unlearned tail. We prove that reducible loss is asymptotically determined by the probability mass of the tail a resource-dependent frontier truncation. Based on our framework, we derive the precise scaling laws for $N$, $D$, and $C$, attributing them to capacity, coverage, and optimization bottlenecks, respectively. Furthermore, we unify these mechanisms via a Max-Bottleneck principle, demonstrating that the Kaplan and Chinchilla scaling laws are not contradictory, but equilibrium solutions to the same constrained optimization problem under different active bottlenecks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。