通过计算规模定律提升高能物理喷注分类性能。
Neural Scaling Laws for Boosted Jet Tagging
- 基于大规模数据与模型扩展研究神经网络的缩放规律。
- 发现增加计算资源可稳定逼近性能上限,提升分类准确率。
- 适合关注高能物理中机器学习模型优化的研究者。
大型语言模型的成功表明,通过同时增加模型容量和数据量来扩大计算规模,是现代机器学习性能提升的主要驱动力。尽管机器学习早已融入高能物理(HEP)数据分析流程,但当前最先进的HEP模型所用计算量仍远低于产业界基础模型。随着该领域缩放规律研究刚刚起步,本文使用公开的JetClass数据集,研究了强化喷注分类中的神经缩放规律。我们推导出计算最优缩放规律,识别出可通过增加计算量持续逼近的有效性能极限。研究发现,在高能物理中因仿真成本高而常见的数据重复行为会改变缩放关系,带来可量化的有效数据规模增益。此外,我们分析了输入特征和粒子多重性对缩放系数及渐近性能极限的影响,结果表明:更多计算量可稳定推动性能趋近极限;更丰富、低层级的特征能提高性能上限,并在固定数据规模下改善结果。
原文摘要 · Abstract (English)
The success of Large Language Models (LLMs) has established that scaling compute, through joint increases in model capacity and dataset size, is the primary driver of performance in modern machine learning. While machine learning has long been an integral component of High Energy Physics (HEP) data analysis workflows, the compute used to train state-of-the-art HEP models remains orders of magnitude below that of industry foundation models. With scaling laws only beginning to be studied in the field, we investigate neural scaling laws for boosted jet classification using the public JetClass dataset. We derive compute optimal scaling laws and identify an effective performance limit that can be consistently approached through increased compute. We study how data repetition, common in HEP where simulation is expensive, modifies the scaling yielding a quantifiable effective dataset size gain. We then study how the scaling coefficients and asymptotic performance limits vary with the choice of input features and particle multiplicity, demonstrating that increased compute reliably drives performance toward an asymptotic limit, and that more expressive, lower-level features can raise the performance limit and improve results at fixed dataset size.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。