用可验证方法将异常检测得分转为可靠显著性,解决高能物理新物理搜寻中的误报问题。
Conformal calibration and look-elsewhere effect in anomaly detection for new-physics searches
- 基于共形预测构建校准层,无需重训练即可纠正得分偏差。
- 在真实对撞机数据上,将虚假46σ信号降为零,避免背景扰动导致的误判。
- 适合新物理搜索中需严格控制假阳性、追求可审计性的研究者使用。
机器学习异常检测正重塑新物理搜寻,但其统计解释滞后。原始异常得分无校准意义,多区域扫描放大‘他处寻找’效应,且依赖的渐近显著性对异常检测器易受的背景建模偏差视而不见。本文提出基于共形预测的校准层,将任意异常得分转化为具有分布无关、有限样本保证的可辩护显著性。共形预测将得分转为有效局部p值,加权与莫德里安变体修复共振搜寻中边带与信号区交换性失效问题,Gross-Vitells步骤则推导出考虑试错因子的全局显著性。该层同时暴露并修正标准流程无法察觉的校准偏差,无需重新训练探测器。在公开的LHC奥运会数据上,分类器产生子结构-质量相关性,使边带校准的背景p值严重偏小;若不加修正,仅由背景雕刻即可制造约46σ过剩。加权校正后恢复真实零假设。盲态宽质量范围搜寻中,标准渐近与未加权方法在无信号窗口仍虚构出>10σ和≈5σ过剩,而本方法未产生任何假警报,其全局误报率经纯背景伪实验验证。结果提供了一条从非校准得分到试错因子感知显著性的可审计、探测器无关路径,可直接融入实验异常搜寻流程。
原文摘要 · Abstract (English)
Machine-learned anomaly detection is reshaping searches for new physics, but it has outrun the statistics used to interpret it. A raw anomaly score has no calibrated meaning, a model that scans many regions inflates the look-elsewhere effect, and the asymptotic significances the field relies on are blind to the background mismodelling that anomaly detectors are especially prone to. We propose a calibration layer, built on conformal prediction, that turns any anomaly score into a defensible significance with distribution-free, finite-sample guarantees. Conformal prediction converts scores into valid local p-values, weighted and Mondrian variants repair the sideband-to-signal-region exchangeability failures that resonant searches suffer, and a Gross-Vitells step carries the result through to a look-elsewhere-aware global significance. The layer does two things at once. It exposes miscalibration that the standard pipeline cannot see, and it corrects it without retraining the detector. On public LHC Olympics data, a classifier develops a substructure-mass correlation that makes sideband-calibrated background p-values anti-conservative. Taken at face value, this manufactures a $\sim 46σ$ excess from background sculpting alone, which the label-free weighted correction removes, restoring an honest null. When run as a blind wide-mass bump hunt, the standard asymptotic and unweighted procedures fabricate $\gtrsim10σ$ excesses and $\approx5σ$ excesses even in signal-free windows, while the conformal layer raises no false alarms and its global false-positive rate is verified on background-only pseudoexperiments. The result is an auditable, detector-agnostic path from an uncalibrated score to a trials-factor-aware significance, ready to be folded into experimental anomaly searches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。