改进粒子生成预训练,提升模型分类能力。
Enhancing next token prediction based pre-training for jet foundation models
- 用连续特征向量输入,保留令牌ID作为预测目标。
- 结合掩码粒子建模与生成任务,提升下游分类效果。
- 不损失生成性能,适合需要高精度分类的粒子研究。
下一令牌预测是构建喷注基础模型的一种有吸引力的预训练任务,因其无需模拟且能实现出色的生成能力并跨数据集迁移。本文研究了多项对下一令牌预测的改进,基于OmniJet-α的初步工作。不同于将粒子分词后仅使用令牌ID作为生成和分类任务的输入,我们采用混合设置,允许使用连续特征向量作为模型输入,而仅在下一令牌预测的目标中使用令牌ID。其次,我们探索了一种结合掩码粒子建模和生成学习目标的联合预训练策略。综合这些改进,在不损失生成性能的前提下,显著提升了下游分类任务的表现。
原文摘要 · Abstract (English)
Next token prediction is an attractive pre-training task for jet foundation models, in that it is simulation free and enables excellent generative capabilities that can transfer across datasets. Here we study multiple improvements to next token prediction, building on the initial work of OmniJet-$α$. Instead of tokenizing particles and subsequently only using the token-ID as the model input for both the generative and the classification task, we adopt a hybrid setup, which allows us to use continuous feature vectors as model input while only using token-IDs in the next token prediction target. Secondly, we explore a combined pre-training strategy that combines masked particle modeling and generative learning objectives. Taken together, these changes greatly improve the performance in downstream classification tasks without any loss in generative performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。