用小模型预训练数据预测大模型性能,提前规划粒子物理模型的算力投入。
Predict before you train: Scaling Laws for particle physics foundation models

- 基于小规模模型拟合联合模型-数据缩放定律,实现对大规模模型的精准性能预测。
- 预测误差低于1%,且预训练损失越低,下游任务的背景抑制能力越强。
- 适合需要高效分配算力的粒子物理研究者,尤其关注模型训练成本与性能平衡者。
粒子物理中最大的机器学习模型训练成本极高,但其扩展性能在投入算力前难以评估。尽管已有针对喷注(jets)的缩放定律,但尚未有方法能预测未参与拟合的模型表现。本文证明,对于在对撞机喷注上预训练的通用Transformer模型,其性能可被准确预测。仅通过三个数量级训练算力范围内小模型的联合模型-数据缩放定律拟合,即可将后续训练超过百倍算力的大模型损失预测误差控制在1%以内。进一步表明:预训练损失越低,微调后损失越小,背景拒识率越高。在两个标准标签基准上,该方法可将算力预算转化为预期物理性能。最终模型在准确率、AUC及夸克/胶子区分度方面与当前先进物理感知基础模型一致,仅在顶部标签高纯度尾部略有优势。本文发布五个不同规模的预训练模型,以及完整的训练配方与代码。
原文摘要 · Abstract (English)
The largest machine learning models in particle physics are also the most expensive to train, yet the return on scaling a given architecture cannot be estimated before that compute is spent. Scaling laws have been fit for jets, but none has yet been shown to predict the performance of models it was not fit on. We show that, for a generic transformer pretrained on collider jets, it can be forecast. Fitting a joint model-and-data scaling law on small models alone, spanning three orders of magnitude of training compute, we predict the loss of models trained afterward with more than one hundred times more compute to within one percent. We then connect the forecast to downstream physics performance: across two standard tagging benchmarks, lower pretraining loss yields systematically lower fine-tuning loss and higher background rejection after fine-tuning. Within this model family and these tasks, a compute budget can therefore be translated into expected physics performance before any large model is trained. The final frontier model is consistent with the published numbers for current state-of-the-art physics-aware foundation models trained on the same corpus, on accuracy, AUC, and quark/gluon rejection, with a residual edge for the physics-aware model only in the high-purity tail of top tagging. We release five pretrained models spanning multiple sizes, together with the complete training recipe and code.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。