arXiv:2506.14934cs.CV2025-06中稿 · Third Internationa…被引 2

用视觉Transformer提升粒子对撞机喷注分类精度

Vision Transformers for End-to-End Quark-Gluon Jet Classification from Calorimeter Images

  • 直接处理探测器图像,结合多通道能量沉积数据端到端学习
  • ViT混合模型在F1、AUC等指标上超越传统CNN基线
  • 首次提供公开数据集与可复现的高性能基准框架

区分夸克与胶子发起的喷注是高能物理中关键且具挑战性的任务,对大型强子对撞机的新物理搜索和精密测量至关重要。尽管深度学习(尤其是卷积神经网络)已在基于图像表示的喷注标记中取得进展,但以捕捉全局上下文信息著称的视觉变压器(ViT)架构在真实探测器与堆叠条件下的直接应用仍研究不足。本文系统评估了ViT及ViT-CNN混合模型在2012年CMS开放数据上的夸克-胶子喷注分类表现。通过构建包含电磁量能器(ECAL)、强子量能器(HCAL)和重建轨迹的多通道喷注视图图像,实现了端到端学习。全面基准测试表明,基于ViT的模型(尤其是ViT+MaxViT和ViT+ConvNeXt混合模型)在F1分数、ROC-AUC和准确率上持续优于现有CNN基线,凸显其对喷注内长程空间关联的建模优势。本工作建立了首个系统性框架与稳健性能基准,为使用公开对撞机数据开展基于图像的喷注分类深度学习研究提供了支持。

原文摘要 · Abstract (English)

Distinguishing between quark- and gluon-initiated jets is a critical and challenging task in high-energy physics, pivotal for improving new physics searches and precision measurements at the Large Hadron Collider. While deep learning, particularly Convolutional Neural Networks (CNNs), has advanced jet tagging using image-based representations, the potential of Vision Transformer (ViT) architectures, renowned for modeling global contextual information, remains largely underexplored for direct calorimeter image analysis, especially under realistic detector and pileup conditions. This paper presents a systematic evaluation of ViTs and ViT-CNN hybrid models for quark-gluon jet classification using simulated 2012 CMS Open Data. We construct multi-channel jet-view images from detector-level energy deposits (ECAL, HCAL) and reconstructed tracks, enabling an end-to-end learning approach. Our comprehensive benchmarking demonstrates that ViT-based models, notably ViT+MaxViT and ViT+ConvNeXt hybrids, consistently outperform established CNN baselines in F1-score, ROC-AUC, and accuracy, highlighting the advantage of capturing long-range spatial correlations within jet substructure. This work establishes the first systematic framework and robust performance baselines for applying ViT architectures to calorimeter image-based jet classification using public collider data, alongside a structured dataset suitable for further deep learning research in this domain.

粒子物理视觉Transformer喷注分类深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。