arXiv:2605.14567stat.MLcs.LG2026-05

揭示深度网络中特征逐级学习如何催生平滑的缩放定律

Scaling Laws from Sequential Feature Recovery: A Solvable Hierarchical Model

  • 基于分层特征结构设计分层谱算法,实现渐进式特征恢复
  • 弱特征需更多数据才能识别,强特征在小样本下即可见
  • 理论证明特征恢复阈值,解释误差随数据量呈幂律下降

我们提出一种简单机制,解释多层网络中特征学习如何催生缩放定律。研究一个高维分层目标:全局高阶函数,但可由一组权重按幂律衰减的潜在组合特征表示。我们发现,适配此组合结构的逐层谱算法相比浅层非自适应方法具有更优缩放性能,并能顺序恢复潜在方向:强特征在小样本下即可检测,弱特征则需更多数据。我们证明了特征级别的精确恢复阈值,表明这些转变的累积导致预测误差显式呈现幂律衰减。技术上,分析依赖随机矩阵方法和基于预解算子的扰动论证,给出了超越标准间距扰动界、对单个特征向量恢复的匹配上下界。数值实验验证了预测的逐级恢复现象、有限样本下的阈值平滑及与非分层核基线的分离。结果共同说明,平滑缩放定律可源于一系列尖锐的特征学习转变。

原文摘要 · Abstract (English)

We propose a simple mechanism by which scaling laws emerge from feature learning in multi-layer networks. We study a high-dimensional hierarchical target that is a globally high-degree function, but that can be represented by a combination of latent compositional features whose weights decrease as a power law. We show that a layer-wise spectral algorithm adapted to this compositional structure achieves improved scaling relative to shallow, non-adaptive methods, and recovers the latent directions sequentially: strong features become detectable at small sample sizes, while weaker features require more data. We prove sharp feature-wise recovery thresholds and show that aggregating these transitions yields an explicit power-law decay of the prediction error. Technically, the analysis relies on random matrix methods and a resolvent-based perturbation argument, which gives matching upper and lower bounds for individual eigenvector recovery beyond what standard gap-based perturbation bounds provide. Numerical experiments confirm the predicted sequential recovery, finite-size smoothing of the thresholds, and separation from non-hierarchical kernel baselines. Together, these results show how smooth scaling laws can emerge from a cascade of sharp feature-learning transitions.

缩放定律特征学习分层模型随机矩阵

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。