arXiv:2511.15779hep-phcs.LG2025-11被引 4

用生成模型验证喷注分类器极限,发现现有方法已逼近理论上限。

SURFing to the Fundamental Limit of Jet Tagging

  • 构建SURF框架,通过可计算的代理模型验证生成模型有效性。
  • 实证表明现代喷注分类器性能接近统计理论极限。
  • 揭示GPT类模型夸大区分能力,可能误导对极限的认知。

除了提升搜索与测量灵敏度的实际目标外,一个更深层的问题是:喷注分类算法的性能上限是什么?具有学习似然函数的生成代理模型为此提供了新路径,前提是该代理能准确捕捉真实数据分布。本文提出SUrrogate ReFerence(SURF)方法,用于验证生成模型的有效性。该框架通过在另一个可计算的代理模型上训练目标模型,实现精确的Neyman-Pearson检验,而该代理模型本身由真实数据训练得到。我们论证EPiC-FM生成模型可作为JetClass喷注的合理代理参考,并应用SURF表明现代喷注分类器可能已接近真正的统计极限。相比之下,自回归GPT模型过度夸大了代理参考中的顶夸克与胶子喷注分离能力,暗示其对根本极限的判断存在偏差。

原文摘要 · Abstract (English)

Beyond the practical goal of improving search and measurement sensitivity through better jet tagging algorithms, there is a deeper question: what are their upper performance limits? Generative surrogate models with learned likelihood functions offer a new approach to this problem, provided the surrogate correctly captures the underlying data distribution. In this work, we introduce the SUrrogate ReFerence (SURF) method, a new approach to validating generative models. This framework enables exact Neyman-Pearson tests by training the target model on samples from another tractable surrogate, which is itself trained on real data. We argue that the EPiC-FM generative model is a valid surrogate reference for JetClass jets and apply SURF to show that modern jet taggers may already be operating close to the true statistical limit. By contrast, we find that autoregressive GPT models unphysically exaggerate top vs. QCD separation power encoded in the surrogate reference, implying that they are giving a misleading picture of the fundamental limit.

喷注分类生成模型统计极限机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。