arXiv:2606.19781hep-excs.AI2026-06

通过设计预训练数据,让物理模型更依赖数据而非增大模型规模。

Towards Engineering Scaling Laws with Pretraining Data Composition

论文配图:Towards Engineering Scaling Laws with Pretraining Data Composition
图 1 · 摘自论文原文
  • 用多样化且任务相关的合成数据优化预训练集
  • 使性能提升更依赖数据量而非模型参数量
  • 适合需要高效扩展的高能物理模型开发

神经网络的缩放定律描述了模型性能随计算量、模型规模和数据规模呈幂律增长。尽管在大语言模型中已得到充分验证,这类关系在粒子物理的大模型中仍处于探索阶段。与自然语言或图像领域不同,高能物理拥有高保真模拟器,可低成本生成合成数据,这使得增加数据比增加参数更经济。我们针对高能粒子对撞中产生的强子喷注分类任务发现,通过引入与下游任务更匹配、更具多样性的预训练数据,可主动调控缩放行为,使模型性能提升更依赖数据量而非模型规模,从而实现以数据驱动的高效扩展。

原文摘要 · Abstract (English)

Neural scaling laws describe how model performance improves as a power law in compute, model size, and dataset size. While well-established for large language models, these relationships are emerging for large models in particle physics. As with language, empirical studies show that the performance scales as a power law. However, unlike natural language or image domains, fundamental physics has high-fidelity simulators that produce synthetic data cheaply. This favors scaling regimes where additional data is cheaper than additional parameters, and allows the pretraining dataset itself to be engineered to influence the scaling. For the task of classifying hadronic jets produced in collisions of high-energy particle beams, we show that the scaling behavior can be engineered towards requiring more data rather than larger models by inclusion of pretraining data which is more diverse and better aligned with the downstream classification task.

缩放定律粒子物理数据工程合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。