arXiv:2606.14870hep-phcs.LG2026-06

对比三种预训练方法,发现生成与分类任务需分别预训练才能有效迁移。

Pre-Training for Simulation-Based Science: A Study on Jet Foundation Model Training Objectives

论文配图:Pre-Training for Simulation-Based Science: A Study on Jet Foundation Model Training Objectives
图 1 · 摘自论文原文
  • 用标注数据、生成建模和掩码粒子建模三种方式预训练模型。
  • 标签少时,分类+掩码建模效果最好;生成任务必须用流匹配预训练。
  • 揭示分类与生成任务的独立性,指导科学领域模型设计。

基于大规模仿真数据的科学基础模型(FMs)在人工智能赋能科学中展现出强大潜力。工业界常用自监督掩码方法,而科学领域因仿真精确且数据丰富,可构建大规模有标签数据集,为预训练提供新路径。本文在OmniLearned高能物理基础模型框架下,系统比较了监督分类、流匹配生成和自监督掩码粒子建模(MPM)三种预训练策略。所有模型均在JetClass数据集上预训练,并在顶喷注分类与JetNet条件生成两个下游任务上微调。结果表明:当下游标签充足时,纯分类预训练最优;但标签稀缺时,结合分类与掩码粒子建模(MPM)效果最佳。流匹配生成预训练对分类任务提升有限,而生成任务仅在预训练中包含流匹配时才能显著受益,暗示分类与生成任务具有正交性——模型要同时迁移到两类任务,必须在预训练阶段覆盖两者。该研究为仿真驱动科学中的基础模型预训练目标提供了可控的规模化分析范式。

原文摘要 · Abstract (English)

Foundation models (FMs) trained on large datasets and fine-tuned on downstream tasks have emerged as a powerful paradigm in AI for science. Industrial FMs are typically trained using self-supervision with masking due to the lack of labels. In many scientific domains, accurate simulations are plentiful and facilitate large, labeled datasets. This opens up new possibilities for pre-training. We present a systematic comparison of pre-training methods using the OmniLearned High Energy Physics FM framework. We test supervised classification, flow-matching generation, and self-supervised masked particle modeling. All models are pre-trained on the JetClass dataset and fine-tuned on two representative downstream tasks, top jet classification and JetNet conditional generation. Among other observations, for classification tasks, we find that pure classifier pre-training is optimal when downstream labels and model capacity are plentiful, but combining it with self-supervised masked particle modeling (MPM) is uniquely powerful in the low-finetuning label regime. Flow matching-based generative pre-training seems to provide little benefit for downstream classification, and interestingly, for downstream generation, we find that flow matching must be in the pre-training objective to see a significant finetuning advantage, hinting at the orthogonality of classification and generation tasks. That is, for a model to transfer to both generative and classification downstream tasks, it must be pre-trained on both. This study provides a template for controlled scaling analysis of pre-training objectives for foundation models in simulation-based sciences.

基础模型高能物理预训练生成建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。