arXiv:2602.10300cs.LG2026-02被引 4

用大模型学习训练配置与性能的关系,提升大规模训练预测精度。

Configuration-to-Performance Scaling Law with Neural Ansatz

  • 用大语言模型拟合完整训练配置到性能的映射关系
  • 预测误差比经典定律低20%-40%,支持10倍以上算力外推
  • 可联合调参,且能扩展到损失曲线预测,适合大规模实验设计

研究人员构建缩放定律以预测大规模训练在更大模型规模N和数据规模D下的性能。这些定律假设其他训练超参数已最优选择,但这一过程可能耗时巨大,甚至因外部硬件限制而无法实现。为提升在更广泛超参数范围内的可预测性,并简化大规模调参,我们提出学习一种 extit{配置到性能的缩放定律}(CPL):从完整的训练配置到训练性能的映射。由于该映射无简单函数形式,我们采用大语言模型(LLM)参数化,并基于多个来源的开源预训练日志进行拟合,得到 extit{神经型}配置到性能缩放定律(NCPL)。NCPL能准确预测训练配置对最终预训练损失的影响,预测误差比无配置依赖的Chinchilla定律低20%-40%,并能推广至训练计算量高达训练集10倍的场景。它还支持多超参数联合调优,性能媲美超参数缩放定律基线。此外,NCPL自然且有效地扩展至更丰富的预测目标,如损失曲线预测。

原文摘要 · Abstract (English)

Researchers build scaling laws to forecast the training performance of expensive large-scale runs with larger model size N and data size D. These laws assume that other training hyperparameters are optimally chosen, which can require significant effort and, in some cases, be impossible due to external hardware constraints. To improve predictability across a broader set of hyperparameters and enable simpler tuning at scale, we propose learning a \textit{Configuration-to-Performance Scaling Law} (CPL): a mapping from the \textit{full training configuration} to training performance. Because no simple functional form can express this mapping, we parameterize it with a large language model (LLM), and fit it with diverse open-source pretraining logs across multiple sources, yielding a \textit{Neural} Configuration-to-Performance Scaling Law (NCPL). NCPL accurately predicts how training configurations influence the final pretraining loss, achieving 20-40% lower prediction error than the configuration-agnostic Chinchilla law and generalizing to runs using up to 10 x more compute than any run in the training set. It further supports joint tuning of multiple hyperparameters with performance comparable to hyperparameter scaling law baselines. Finally, NCPL naturally and effectively extends to richer prediction targets such as loss-curve prediction.

缩放定律大模型训练优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。