揭示大模型知识迁移中的特征学习突变与最优泛化规律
Sharp feature-learning transitions and Bayes-optimal neural scaling laws in extensive-width networks
- 通过理论推导发现特征可学性存在一系列突变跃迁
- 数据量增长时特征逐个被恢复,出现重叠率的不连续跳跃
- 提出有效宽度概念,统一解释不同数据规模下的最优泛化
我们研究在高维极限下,从噪声查询中学习具有分层特征的一层教师网络的信息论极限,目标是将知识转移给更小的学生模型。考虑教师宽度k随输入维度d线性增长的情形,该设置适用于大但有限宽的神经网络,且近期才具备解析可行性。基于启发式的留一法解耦论证(经数值验证),我们推导出贝叶斯最优泛化误差和单个特征重叠的渐近精确表征,其由一组封闭的不动点方程描述。这些方程揭示:特征可学习性受一系列尖锐相变支配——随着数据量n增加,教师特征按序被恢复,每次通过重叠率的不连续跳跃实现。这一序列获取机制定义了精确的‘有效宽度’k_c,即在给定数据预算n下可学习特征的数量。它统一了两种不同标度律:在特征学习阶段,贝叶斯最优泛化误差ε^BO ∝ n^(1/(2β)-1);在精炼阶段,ε^BO ∝ n^(-1),其中β>1/2为幂律特征层级的指数。两种标度律统一为ε^BO = Θ(k_c d / n)。我们进一步实证表明,以Adam训练的学生在接近有效宽度k_c时,可达到这些最优标度律(仅存在微小算法间隙),并提供了模型规模相关标度的信息论解释。
原文摘要 · Abstract (English)
We study the information-theoretic limits of learning a one-hidden-layer teacher network with hierarchical features from noisy queries, in the context of knowledge transfer to a smaller student model. We work in the high-dimensional regime where the teacher width $k$ scales linearly with the input dimension $d$ -- a setting that captures large-but-finite-width networks and has only recently become analytically tractable. Using a heuristic leave-one-out decoupling argument, validated numerically throughout, we derive asymptotically sharp characterizations of the Bayes-optimal generalization error and individual feature overlaps via a system of closed fixed-point equations. These equations reveal that feature learnability is governed by a sequence of sharp phase transitions: as data grows, teacher features become recoverable sequentially, each through a discontinuous jump in overlap. This sequential acquisition underlies a precise notion of \textit{effective width} $k_c$ -- the number of learnable features at a given data budget $n$ -- which unifies two distinct scaling regimes: a feature-learning regime in which the Bayes-optimal generalization error $\varepsilon^{\rm BO}$ scales as $ n^{1/(2β)-1}$, and a refinement regime in which it scales as $n^{-1}$, where $β>1/2$ is the exponent of the power-law feature hierarchy. Both laws collapse to the single relation $\varepsilon^{\rm BO}=Θ(k_c d/n)$. We further show empirically that a student trained with \textsc{Adam} near the effective width $k_c$ achieves these optimal scaling laws (up to a small algorithmic gap), and provide an information-theoretic account of the associated scaling in model size.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。