arXiv:2605.28940hep-phcs.LG2026-05被引 2

首次发现喷注生成模型存在参数规模的对数缩放规律,验证了损失函数可反映物理性能。

Neural Scaling Laws for Jet Generation

  • 通过自回归预测喷注粒子,研究模型参数、数据量和算力的缩放关系
  • 模型参数增大时,验证损失与物理量距离均呈对数下降,但数据和算力缩放弱
  • 提出可学习窗口概念,解释喷注生成饱和快于语言模型的原因

最近观察到的经验缩放规律描述了基础模型性能随数据量、算力和模型参数变化的关系。本研究首次探索喷注生成任务是否存在此类规律,该任务既可用于基础模型预训练,也可作为原位模拟。我们确实在模型规模上复现了关键的对数缩放行为。除了分析生成模型的下一个词预测验证损失外,还研究了五个训练中不可见的物理量的切片沃尔什距离。结果表明该距离与验证损失单调相关,说明损失是物理性能的良好代理。对于数据量和算力的缩放,损失和切片沃尔什距离的缩放行为明显较弱。我们引入可学习窗口概念,认为喷注粒子的自回归预测相比语言模型更快饱和。可能原因包括量子色动力学辐射的随机性,以及生成与监督学习在对撞机物理中的差异。

原文摘要 · Abstract (English)

Recently observed empirical scaling laws describe the performance of foundation-type models as three independent key quantities -- dataset size, compute, and model parameters -- are modified. Extracting these scaling laws informs the training of large complex models for which the tuning of hyperparameters in traditional ways is not feasible. This work for the first time explores if scaling laws can also be observed for the task of particle jet generation -- both relevant as a pre-training objective for foundation models and as in-situ simulation by itself. We indeed replicate the key logarithmic scaling law behavior for model-size scaling. Beyond studying the next token prediction validation loss of the generative model, we also study the sliced Wasserstein distance of five physical quantities that are not immediately available to the model during training. Our study shows that this quantity is monotonically related to the next token prediction validation loss, meaning that this loss is indeed a good proxy for the physics performance. For the scaling with dataset size and compute, we observe substantially weaker scaling behavior of both the loss and the sliced Wasserstein distance. We analyze this behavior by introducing the concept of a learnable window, and argue that autoregressive next token prediction on jet constituents exhibits comparatively rapid saturation relative to language-model studies. We discuss possible origins of this behavior, including the stochastic nature of QCD radiation and differences between generative and supervised learning tasks in collider physics.

喷注生成缩放定律生成模型物理模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。