arXiv:2409.12589hep-phcs.LG2024-09被引 27

无需分词也能高效建模粒子集,提升高能物理基础模型性能

Is Tokenization Needed for Masked Particle Modelling?

  • 不依赖分词的条件生成重建方法,直接处理原始粒子数据
  • 在喷注任务上超越原有分词方法,多个下游任务表现更优
  • 适合高能物理领域研究者构建无标签基础模型

本文显著提升了面向无序粒子集的掩码粒子建模(MPM)方法,这是一种用于构建高能物理领域基础模型的自监督学习框架。在该框架中,模型通过恢复集合中缺失元素来训练,无需标签,可直接应用于实验数据。我们通过优化实现效率并引入更强的解码器,大幅提升了性能。对比多种预训练任务,提出基于条件生成模型的新重建方法,避免了数据分词或离散化。在包含分类、次级顶点寻找和轨迹识别等多样化下游任务的新测试基准上,新方法优于原始MPM的分词学习目标。

原文摘要 · Abstract (English)

In this work, we significantly enhance masked particle modeling (MPM), a self-supervised learning scheme for constructing highly expressive representations of unordered sets relevant to developing foundation models for high-energy physics. In MPM, a model is trained to recover the missing elements of a set, a learning objective that requires no labels and can be applied directly to experimental data. We achieve significant performance improvements over previous work on MPM by addressing inefficiencies in the implementation and incorporating a more powerful decoder. We compare several pre-training tasks and introduce new reconstruction methods that utilize conditional generative models without data tokenization or discretization. We show that these new methods outperform the tokenized learning objective from the original MPM on a new test bed for foundation models for jets, which includes using a wide variety of downstream tasks relevant to jet physics, such as classification, secondary vertex finding, and track identification.

自监督学习高能物理粒子建模无分词

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。