稀疏激活导致模型性能出现双下降,且算力最优配置更倾向扩大数据集。
Asymmetric Scaling Laws from Sparse Features
- 基于稀疏激活构建新模型,揭示测试误差由训练中未出现的罕见特征主导。
- 在参数量临界点附近出现双下降峰值,过参数与欠参数区分别遵循不同幂律。
- 算力有限时应优先扩充数据而非增大模型,适用于稀疏输入场景的模型设计。
我们提出一种神经网络缩放定律的稀疏激活模型。该模型显示,测试损失常由训练数据中从未出现的罕见坐标主导,由此产生稀疏模型特有的瓶颈。我们推导了欠参数和过参数两种情形下的渐近总体损失,发现损失在参数恰好能拟合训练数据的插值阈值附近呈现双下降峰,表现出两个不同的缩放指数,其差距由稀疏度决定。此外,我们推导出固定算力预算下的计算最优前沿,表明应优先增加数据规模而非模型容量。我们还分析了梯度下降动力学,给出固定步长梯度下降失稳概率的缩放定律,并证明稀疏性效应在非线性激活下依然存在。
原文摘要 · Abstract (English)
We introduce a model for neural scaling laws under sparse activations. In the model, test loss is often dominated by rare coordinates that are never observed in the training input. This mechanism induces a novel bottleneck absent from dense models. We derive the asymptotic population loss in both the underparameterized and overparameterized regimes, and show that the loss exhibits a double-descent peak near the interpolation threshold -- where the number of parameters is just sufficient to fit the training data -- resulting in a loss curve governed by two distinct scaling exponents -- one for the overparameterized regime and one for the underparameterized regime -- with a gap determined by the degree of sparsity. Additionally, we derive a compute-optimal frontier that favors increasing dataset size over model capacity under fixed compute budgets. We also analyze gradient-descent dynamics and identify a scaling law for the probability that fixed-step gradient descent becomes unstable. We further show that the sparsity-induced effect persists under nonlinear activations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。