arXiv:2412.07942cs.LGcond-mat.dis-nn2024-12被引 13

用渗流理论解释深度模型的通用缩放规律,揭示数据分布如何决定性能提升方式。

Neural Scaling Laws Rooted in the Data Distribution

  • 基于渗流理论构建数据分布模型,模拟自然任务的结构特性。
  • 发现两种临界状态,分别对应幂律子任务与主导数据流形,均产生最优缩放律。
  • 在模拟数据上验证理论,为语言模型缩放预测提供新方向。

深度神经网络在多种架构、任务和数据集上均表现出经验性的神经缩放规律:误差随模型或数据规模增大而按幂律下降。这种普适性表明缩放规律可能源于自然学习任务的普遍性质。我们提出一个数学模型,利用渗流理论描述自然数据分布。模型中出现两种不同的临界态,各自导出最优的幂律缩放规律。这两种状态分别对应幂律分布的离散子任务和主导数据流形,可与此前提出的缩放理论关联,从而统一并扎根于已有工作。通过在基于渗流模拟生成的合成数据集上训练回归模型,我们验证了该理论。研究还提出了定量预测语言模型缩放行为的方向。

原文摘要 · Abstract (English)

Deep neural networks exhibit empirical neural scaling laws, with error decreasing as a power law with increasing model or data size, across a wide variety of architectures, tasks, and datasets. This universality suggests that scaling laws may result from general properties of natural learning tasks. We develop a mathematical model intended to describe natural datasets using percolation theory. Two distinct criticality regimes emerge, each yielding optimal power-law neural scaling laws. These regimes, corresponding to power-law-distributed discrete subtasks and a dominant data manifold, can be associated with previously proposed theories of neural scaling, thereby grounding and unifying prior works. We test the theory by training regression models on toy datasets derived from percolation theory simulations. We suggest directions for quantitatively predicting language model scaling.

神经网络缩放定律数据分布渗流理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。