用LHC真实数据训练粒子物理基础模型,提升生成新喷注能力
Aspen Open Jets: Unlocking LHC Data for Foundation Models in Particle Physics
- 基于CMS实验2016年开放数据构建1.78亿个高动量喷注数据集
- 在真实喷注数据上预训练后,生成性能显著优于仅用模拟数据训练
- 开源数据集支持后续研究,推动粒子物理领域基础模型发展
基础模型是通过大规模数据预训练的深度学习模型,具备跨数据集和下游任务的泛化能力。本文展示了大型强子对撞机(LHC)CMS实验收集的数据如何用于粒子物理中基础模型的预训练。具体而言,我们提出了AspenOpenJets数据集,包含约1.78亿个来自CMS 2016年开放数据的高 $p_T$ 喷注。我们证明,在AspenOpenJets上预训练OmniJet-$α$基础模型,可显著提升其在存在显著领域偏移情况下的生成性能:从模拟的JetClass数据集中生成增强型顶夸克与QCD喷注。本工作不仅验证了基于真实质子-质子碰撞数据预训练喷注基础模型的有效性,还提供了可直接用于机器学习的AspenOpenJets数据集供公众使用。
原文摘要 · Abstract (English)
Foundation models are deep learning models pre-trained on large amounts of data which are capable of generalizing to multiple datasets and/or downstream tasks. This work demonstrates how data collected by the CMS experiment at the Large Hadron Collider can be useful in pre-training foundation models for HEP. Specifically, we introduce the AspenOpenJets dataset, consisting of approximately 178M high $p_T$ jets derived from CMS 2016 Open Data. We show how pre-training the OmniJet-$α$ foundation model on AspenOpenJets improves performance on generative tasks with significant domain shift: generating boosted top and QCD jets from the simulated JetClass dataset. In addition to demonstrating the power of pre-training of a jet-based foundation model on actual proton-proton collision data, we provide the ML-ready derived AspenOpenJets dataset for further public use.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。