arXiv:2412.03673hep-phcs.LG2024-12中稿 · the Machine Learni…被引 7

解析粒子变压器的注意力机制,揭示其识别喷注的内在逻辑

Interpreting Transformers for Jet Tagging

  • 通过注意力热图分析,发现每个粒子最多关注另一个粒子
  • 模型对重要粒子和次结构的关注随衰变类型变化,体现传统观测量学习能力
  • 为高能物理中优化变压器架构提供可解释性依据

机器学习算法,尤其是基于注意力的变压器模型,已成为分析大型强子对撞机ATLAS和CMS实验产生的海量数据的关键工具。粒子变压器(ParT)作为当前最先进的模型,利用粒子级注意力提升喷注标记任务性能,该任务对识别质子碰撞产物至关重要。本研究通过分析$η$-$ϕ$平面上的注意力热图与粒子对相关性,发现一种二元注意力模式:每个粒子仅关注至多一个其他粒子。同时,观察到模型对关键粒子和次喷注的关注程度随衰变类型而异,表明其已学会传统喷注次结构观测量。这些发现深化了对模型内部运作机制和学习过程的理解,为未来高能物理应用中改进变压器架构效率提供了潜在路径。

原文摘要 · Abstract (English)

Machine learning (ML) algorithms, particularly attention-based transformer models, have become indispensable for analyzing the vast data generated by particle physics experiments like ATLAS and CMS at the CERN LHC. Particle Transformer (ParT), a state-of-the-art model, leverages particle-level attention to improve jet-tagging tasks, which are critical for identifying particles resulting from proton collisions. This study focuses on interpreting ParT by analyzing attention heat maps and particle-pair correlations on the $η$-$ϕ$ plane, revealing a binary attention pattern where each particle attends to at most one other particle. At the same time, we observe that ParT shows varying focus on important particles and subjets depending on decay, indicating that the model learns traditional jet substructure observables. These insights enhance our understanding of the model's internal workings and learning process, offering potential avenues for improving the efficiency of transformer architectures in future high-energy physics applications.

Transformer喷注标记可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。