用无标签数据训练模型,让其在高能物理中更稳定泛化。
RINO: Renormalization Group Invariance with No Labels
- 基于自监督学习,从真实碰撞数据中提取不变特征。
- 在真实数据预训练后,微调到模拟数据上,识别顶夸克喷注准确率提升。
- 适合需要对抗模拟偏差的高能物理机器学习研究者。
高能物理中监督学习常依赖模拟生成的标注数据,但模拟可能无法准确反映真实碰撞或探测器响应。为缓解领域偏移问题,本文提出RINO(无标签下的重整化群不变性),一种自监督学习方法,可直接在真实碰撞数据上预训练模型,学习对重整化群流尺度不变的嵌入表示。我们使用来自JetClass数据集的量子色动力学(QCD)喷注数据预训练一个基于Transformer的模型,模拟真实实验数据;随后在模拟数据集JetNet上微调,完成顶夸克衰变喷注识别任务。结果表明,相比直接在JetNet上从零开始监督训练,RINO在跨域泛化能力上表现更优,验证了在真实数据上预训练、小规模高质量蒙特卡洛数据上微调的可行性,有助于提升高能物理中机器学习模型的鲁棒性。
原文摘要 · Abstract (English)
A common challenge with supervised machine learning (ML) in high energy physics (HEP) is the reliance on simulations for labeled data, which can often mismodel the underlying collision or detector response. To help mitigate this problem of domain shift, we propose RINO (Renormalization Group Invariance with No Labels), a self-supervised learning approach that can instead pretrain models directly on collision data, learning embeddings invariant to renormalization group flow scales. In this work, we pretrain a transformer-based model on jets originating from quantum chromodynamic (QCD) interactions from the JetClass dataset, emulating real QCD-dominated experimental data, and then finetune on the JetNet dataset -- emulating simulations -- for the task of identifying jets originating from top quark decays. RINO demonstrates improved generalization from the JetNet training data to JetClass data compared to supervised training on JetNet from scratch, demonstrating the potential for RINO pretraining on real collision data followed by fine-tuning on small, high-quality MC datasets, to improve the robustness of ML models in HEP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。