用快速模拟数据训练模型,可显著减少真实模拟所需数据量。
Transfer Learning Across Fast- and Full-Simulation Domains in High-Energy Physics
- 在快速与全量模拟间做迁移学习,提升模型泛化能力。
- 各任务均减少约一半目标数据需求,效果稳定。
- 适合高能物理领域模型复用与数据稀缺场景。
高能物理中的机器学习模型通常基于模拟数据训练,全量模拟计算成本高,快速模拟则提供大量但较不真实的样本。本文在真实LHC环境下系统研究了快速模拟与全量模拟数据间的迁移学习,涵盖信号-背景分类、夸克-胶子喷注判别、缺失横动量重建三个典型任务,采用密集网络、图神经网络和基于Transformer的架构。模型先在类似ATLAS的快速模拟上预训练,再迁移到类似CMS的快速模拟及全量模拟的ATLAS开放数据。所有任务中,预训练模型均优于独立训练基线,且所需目标域数据量平均减少约一半。结果表明,快速模拟可学习到稳健可复用的特征表示,支持将训练好的模型作为科学资产发布,超越大型基础模型。
原文摘要 · Abstract (English)
Machine-learning models in high-energy physics are often trained on simulated data, where fully simulated samples are computationally expensive while fast simulation provides large statistics at reduced realism. In this work, we systematically study transfer learning between fast-simulated and fully simulated datasets in a realistic LHC environment. We consider three representative tasks, signal-background classification, quark-gluon jet tagging, and missing transverse energy reconstruction, using dense neural networks, graph neural networks, and transformer-based architectures. Models are pretrained on ATLAS-like fast simulation and adapted to CMS-like fast simulation and to fully simulated ATLAS Open Data. Across all tasks, pretrained models consistently outperform independently trained baselines and require significantly less target-domain training data, typically reducing the needed statistics by about a factor of two. These results demonstrate that fast simulation can be used to learn robust, reusable representations and motivate publishing trained models as reusable scientific assets beyond large foundation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。