arXiv:2601.11719cs.LGhep-ex2026-01

通过自蒸馏学习喷注语义表示,自动聚类出异常检测所需结构。

jBOT: Semantic Jet Representation Clustering Emerges from Self-Distillation

  • 用粒子级与喷注级双层自蒸馏学习喷注表征
  • 仅用背景喷注预训练即实现语义聚类,可直接用于异常检测
  • 冻结嵌入可零样本检测异常,微调后分类性能优于从零训练的监督模型

自监督学习是基础模型训练中无需标签即可学习特征表示的强大方法,常能捕捉数据中的通用语义,并用于下游任务微调。本文提出jBOT,一种基于自蒸馏的喷注数据预训练方法,适用于来自欧洲核子研究中心大型强子对撞机(LHC)的喷注数据。该方法结合局部粒子级蒸馏与全局喷注级蒸馏,学习支持异常检测与分类等下游任务的喷注表示。我们观察到,在未标注喷注上进行预训练后,表示空间中涌现出语义类别聚类。当仅用背景喷注预训练时,冻结的嵌入表示可通过简单的距离度量实现异常检测;而学习到的嵌入可进一步微调,其分类性能优于从零开始训练的监督模型。

原文摘要 · Abstract (English)

Self-supervised learning, in the context of foundation model training, is a powerful pre-training method for learning feature representations without labels, which often capture generic underlying semantics from the data and can later be fine-tuned for downstream tasks. In this work, we introduce jBOT, a pre-training method based on self-distillation for jet data from the CERN Large Hadron Collider, which combines local particle-level distillation with global jet-level distillation to learn jet representations that support downstream tasks such as anomaly detection and classification. We observe that pre-training on unlabeled jets leads to emergent semantic class clustering in the representation space. The clustering in the frozen embedding, when pre-trained on background jets only, enables anomaly detection via simple distance-based metrics, and the learned embedding can be fine-tuned for classification with improved performance compared to supervised models trained from scratch.

自监督学习喷注分析异常检测自蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。