arXiv:2509.20721cs.LGmath.ST2025-09被引 3

揭示深度学习缩放定律的本质是数据冗余规律。

Scaling Laws are Redundancy Laws

  • 从数据协方差谱的多项式尾部出发,推导出学习曲线的幂律关系。
  • 缩放指数由数据冗余度决定,冗余越高,模型收益越快上升。
  • 适用于多种架构与变换,首次实现理论严格解释缩放现象。

缩放定律是深度学习的核心特征,表现为模型性能随数据量和模型规模增加而呈幂律提升。然而其数学根源,特别是缩放指数仍不明确。本文证明缩放定律本质上是冗余定律。通过核回归分析,我们发现数据协方差谱的多项式尾部导致过剩风险呈现幂律,其指数为 alpha = 2s / (2s + 1/beta),其中 beta 控制谱尾行为,1/beta 衡量冗余程度。这表明学习曲线斜率并非普适,而是依赖于数据冗余;谱越陡峭,规模收益越显著。该定律在有界可逆变换、多模态混合、有限宽度近似及 Transformer 架构(包括线性化 NTK 与特征学习两种情形)中均成立。本工作首次以严格数学方式将缩放定律解释为有限样本下的冗余定律,统一了经验观察与理论基础。

原文摘要 · Abstract (English)

Scaling laws, a defining feature of deep learning, reveal a striking power-law improvement in model performance with increasing dataset and model size. Yet, their mathematical origins, especially the scaling exponent, have remained elusive. In this work, we show that scaling laws can be formally explained as redundancy laws. Using kernel regression, we show that a polynomial tail in the data covariance spectrum yields an excess risk power law with exponent alpha = 2s / (2s + 1/beta), where beta controls the spectral tail and 1/beta measures redundancy. This reveals that the learning curve's slope is not universal but depends on data redundancy, with steeper spectra accelerating returns to scale. We establish the law's universality across boundedly invertible transformations, multi-modal mixtures, finite-width approximations, and Transformer architectures in both linearized (NTK) and feature-learning regimes. This work delivers the first rigorous mathematical explanation of scaling laws as finite-sample redundancy laws, unifying empirical observations with theoretical foundations.

缩放定律冗余理论分析深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。