arXiv:2602.00141physics.data-ancs.CV2026-02

用模拟喷注图像对比模型,发现微调Swin-Tiny效果最佳

Comparison of Image Processing Models in Quark Gluon Jet Classification

  • 采用三通道粒子动量表示喷注结构,测试CNN、ViT与Swin-Tiny
  • 微调仅最后两层Swin-Tiny达81.4%准确率,AUC为88.9%
  • 自监督预训练提升特征鲁棒性,适合真实碰撞数据迁移

我们对卷积神经网络(CNN)、视觉变换器(ViTs)和Swin-Tiny模型在区分夸克与胶子喷注方面的表现进行了全面比较。通过将喷注次结构编码为三个通道的粒子动量表示,我们在监督与自监督学习设置下评估了这些模型。结果表明,仅微调Swin-Tiny模型的最后两个块,在效率与精度间取得最佳平衡,准确率达到81.4%,受试者工作特征曲线下面积(AUC)为88.9%。使用动量对比(MoCo)进行自监督预训练进一步增强了特征鲁棒性,并减少了可训练参数数量。这些发现凸显了分层注意力模型在喷注次结构研究中的潜力,以及向真实碰撞数据迁移的应用前景。

原文摘要 · Abstract (English)

We present a comprehensive comparison of convolutional and transformer-based models for distinguishing quark and gluon jets using simulated jet images from Pythia 8. By encoding jet substructure into a three-channel representation of particle kinematics, we evaluate the performance of convolutional neural networks (CNNs), Vision Transformers (ViTs), and Swin Transformers (Swin-Tiny) under both supervised and self-supervised learning setups. Our results show that fine-tuning only the final two transformer blocks of the Swin-Tiny model achieves the best trade-off between efficiency and accuracy, reaching 81.4% accuracy and an AUC (area under the ROC curve) of 88.9%. Self-supervised pretraining with Momentum Contrast (MoCo) further enhances feature robustness and reduces the number of trainable parameters. These findings highlight the potential of hierarchical attention-based models for jet substructure studies and for domain transfer to real collision data.

喷注分类Transformer自监督学习高能物理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。