用代理模型优化计算资源分配,高效估算模型扩展规律。
Active Budget Allocation for Efficient Scaling Law Estimation via Surrogate-Guided Pruning

- 结合代理模型与逐次减半算法,智能分配计算资源。
- 实测在真实和合成数据上分别提升2.84%和5.47%预测精度。
- 相比传统方法节省高达98.7%算力,适合资源受限场景。
预测大规模模型性能有助于设计契合特定目标的训练策略与架构。经验性扩展规律研究通过损失-计算前沿(由学习曲线定义)建立损失与算力之间的函数关系。由于该方法依赖实证,计算开销巨大,因此需战略性分配资源——但这一问题仍鲜被研究。本文探索了逐次减半(SH)算法及其与参数/非参数代理模型结合的适用性。不仅实现更系统的算力分配,还发现结合代理模型的SH能生成低于均匀分配或仅用SH的更低损失-算力组合。实验表明,在真实与合成学习曲线数据集上平均相对提升分别达2.84%和5.47%。该策略显著降低计算成本,相较传统全遍历方法最高可节省98.7%算力。
原文摘要 · Abstract (English)
Predicting model performance at larger scales enables the design of training strategies and architectures tailored to specific performance targets. Empirical scaling law research identifies functional forms to aid this prediction task. These describe the relationship between loss and compute using a loss-compute frontier defined by learning curves. Due to the empirical nature of this approach, the computational burden is substantial, making strategic resource allocation essential - yet it remains surprisingly underexplored. In this work, we address this shortcoming by exploring the suitability of Successive Halving (SH) and SH combined with parametric and non-parametric surrogate models. In addition to enabling a more systematic allocation of a given compute budget, our findings show that SH paired with surrogate models yields a set of learning curves that includes one with a lower loss-compute value than what naive uniform allocation or an SH-only approach can obtain. Our experiments demonstrate mean relative improvements of up to 2.84% and 5.47% on real-world and synthetic learning curve datasets. This strategic resource allocation enables us to obtain accurate scaling laws at significantly reduced computational costs, saving up to 98.7% over the traditional exhaustive approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。