提出新缩放定律,让模型预测更准、省算力。
Skaling: Chinchilla's Exponents Meet Kaplan's Coupling

- 用一个交互指数耦合模型与数据规模,改进传统缩放规律。
- 跨内插和外推场景,误差降低1.5至3倍。
- 小实验即可可靠预测大模型性能,适合资源受限研究者。
神经网络缩放定律是语言模型发展的基础,但标准形式在数据稀缺和过训练极端情况下系统性地低估或高估损失。问题源于假设模型规模与训练数据对损失的影响相互独立。为此,我们提出Skaling定律,一种通过单一交互指数耦合模型容量与数据的广义函数形式。该简单扩展在内插和外推场景中将平均绝对百分比误差(MAPE)降低1.5至3倍。结合稀疏网格策略并限制在低计算范围内,Skaling定律仅需均匀扫描约1/10的计算量即可实现准确的全网格外推。通过使小规模实验即可可靠预测大规模模型性能,该定律为下一代模型训练提供了更鲁棒且资源高效的算力分配框架。
原文摘要 · Abstract (English)
Neural scaling laws are foundational for language model development, yet standard formulations systematically under- and overestimate loss at data-scarce and overtraining extremes. This failure originates in the underlying assumption that model size and training data impact the loss independently. To address this, we introduce the Skaling law, a generalized functional form that couples model capacity and data through a single interaction exponent. This simple extension reduces the Mean Absolute Percentage Error (MAPE) by 1.5-3x across both interpolation and extrapolation regimes. When paired with a sparse grid strategy restricted to low-compute regimes, the Skaling law achieves accurate full-grid extrapolation using approximately 10x less compute than uniform sweeps. By enabling reliable performance prediction from small-scale experiments, the Skaling law provides a more robust and resource-efficient framework for allocating compute budgets in next-generation model training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。