揭示大模型为何更优:能持续学习弱信号。
Spectral Reach: Understanding Neural Scaling as Progress into the Spectral Tail

- 用谱位置衡量模型学习时聚焦的特征谱成分。
- 大模型能深入谱尾,小模型止步于主成分。
- 特征学习是维持弱信号学习的关键,可指导模型设计。
神经网络缩放定律描述了模型规模、数据量、算力与性能间的幂律关系。尽管这些定律指导了现代基础模型的发展,其内在机制仍不清晰,部分源于缺乏可扩展的分析工具。本文提出“谱位置”——一种衡量经验神经正切核(eNTK)哪些特征值主导损失下降的可扩展指标。在缩放实验中发现,训练过程中谱位置持续下降:学习从主特征模式逐步转向谱尾。大模型比小模型能更深地进入谱尾,揭示出一种依赖模型大小的能力,称为“谱可达性”。这解释了为何大模型能取得更低损失:它们可在小模型无法触及的弱谱信号上持续学习。进一步发现,特征学习是实现谱可达性的关键:它随学习推进自适应放大梯度幅值,使模型在固定表示下停滞的区域仍能前进。该发现为架构和优化器设计提供了明确干预方向。
原文摘要 · Abstract (English)
Neural scaling laws describe predictable power-law relationships between model size, dataset size, compute, and performance. While these laws guide the development of modern foundation models, the mechanisms underpinning them remain poorly understood, in part due to the absence of scalable analysis tools. To close this gap, we introduce "spectral position": a scalable measure of which eigenvalues of the empirical neural tangent kernel (eNTK) currently drive loss reduction. Applying this measure to scaling experiments, we find that spectral position decreases throughout training: learning shifts from dominant eigenmodes into the spectral tail. Larger models reach further into the tail than smaller models, revealing a size-dependent capacity we call "spectral reach". This suggests why larger models achieve lower losses: they sustain learning on weak spectral signals inaccessible to smaller models. We further identify feature learning as a key enabler of spectral reach. It adaptively amplifies gradient magnitudes as learning advances, sustaining progress where frozen representations stall. This points to concrete interventions through architecture and optimizer design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。