解析粒子喷注分类模型内部机制,发现关键计算回路及其物理意义。
Dissecting Jet-Tagger Through Mechanistic Interpretability

- 通过零消融与路径修补,定位到六头稀疏计算回路
- 该回路恢复了模型大部分性能,且可解释为源-中继-读出结构
- 模型优先编码二分支子结构,揭示物理信号的隐式学习
本研究对在顶夸克喷注分类基准数据集上训练的Particle Transformer模型进行机制可解释性分析,旨在识别负责喷注分类的计算电路并刻画其内部表征的物理内容。结合零消融、两种互补的流形内扰动路径修补及残差流线性探测,我们识别出一个稀疏的六头回路,该回路恢复了绝大多数全模型性能,并具备清晰的源-中继-读出解释。其中,一个早期层头为主要因果源,一组中间层头作为中继,选择性关注难分辨的成对子结构;一个晚期层头读取聚合信号。线性探测显示,残差流更倾向于对齐能量关联基而非N-亚喷注基。在能量关联基中,模型更侧重编码双分支子结构可观测量而非三分支。逐层探测进一步表明,模型看似在首个类别注意力块即完成分类决策,实为基变换过程,判别信号已在粒子注意力堆栈中饱和。结果表明,自然语言模型的机制可解释方法可用于喷注物理分类器,且梯度下降可能在无监督条件下重新发现物理上有意义的喷注标记特征。
原文摘要 · Abstract (English)
Mechanistic interpretability seeks to reverse engineer a trained neural network by identifying the minimal subset of internal components. We perform a mechanistic interpretability analysis of the Particle Transformer architecture, trained on the Top Quark Tagging reference dataset, with the goal of identifying the computational circuit responsible for jet classification and characterizing the physical content of its internal representations. Combining zero ablation, path patching with two complementary on-manifold corruption strategies and linear probing of the residual stream, we identify a sparse six-head circuit that recovers the great majority of the full model performance while admitting a clean source-relay-readout interpretation. In this circuit, a single early layer head serves as the primary causal source, a cluster of middle-layer heads acts as relays selectively attending to hard pairwise substructure and a single late-layer head reads out the aggregated signal. Linear probes show that the residual stream is preferentially aligned with the energy correlator basis over the $N$-subjettiness basis. Within the energy correlator basis, the model preferentially encodes 2-prong substructure observables over the 3-prong observables. A per-layer trained probe further reveals that the apparent single step commitment of the model to a classification decision in the first class attention block is in fact a basis rotation, with the discriminating signal already saturating in the particle attention stack. These results demonstrate that mechanistic interpretability methods developed for natural language models can be used for jet physics classifiers and indicate that gradient descent may rediscover physically meaningful aspects of jet tagging without supervision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。