用置信预测统一校准高能物理中的机器学习输出,保证误差可控。
Another Fit Bites the Dust: Conformal Prediction as a Calibration Standard for Machine Learning in High-Energy Physics
- 引入置信预测作为通用校准层,无需重训练
- 在多种任务中实现可控的误报率和统计有效性
- 适合需要可靠决策的高能物理实验分析
机器学习在现代对撞机研究中至关重要,但其概率输出常缺乏校准的不确定性估计和有限样本保证,限制了在统计推断与决策中的直接应用。置信预测(Conformal Prediction, CP)提供了一种简单、无需分布假设的校准框架,可在不重新训练模型的情况下为任意预测模型生成严格不确定性量化,并在最小交换性假设下提供有限样本覆盖率保证,无需依赖渐近理论、极限定理或高斯近似。本文研究了CP在高能物理机器学习中的统一校准作用。基于公开对撞机数据集和多种模型,我们表明单一置信形式可适用于回归、二分类、多分类、异常检测和生成建模,将原始模型输出转化为具有控制误报率的预测集、典型性区域和p值。尽管置信预测不提升原始模型性能,但强制实现诚实的不确定性量化与透明的误差控制。我们主张将置信校准作为对撞机物理机器学习流水线的标准组件,以支持可靠解释、稳健比较与原则性统计决策。
原文摘要 · Abstract (English)
Machine-learning techniques are essential in modern collider research, yet their probabilistic outputs often lack calibrated uncertainty estimates and finite-sample guarantees, limiting their direct use in statistical inference and decision-making. Conformal prediction (CP) provides a simple, distribution-free framework for calibrating arbitrary predictive models without retraining, yielding rigorous uncertainty quantification with finite-sample coverage guarantees under minimal exchangeability assumptions, without reliance on asymptotics, limit theorems, or Gaussian approximations. In this work, we investigate CP as a unifying calibration layer for machine-learning applications in high-energy physics. Using publicly available collider datasets and a diverse set of models, we show that a single conformal formalism can be applied across regression, binary and multi-class classification, anomaly detection, and generative modelling, converting raw model outputs into statistically valid prediction sets, typicality regions, and p-values with controlled false-positive rates. While conformal prediction does not improve raw model performance, it enforces honest uncertainty quantification and transparent error control. We argue that conformal calibration should be adopted as a standard component of machine-learning pipelines in collider physics, enabling reliable interpretation, robust comparisons, and principled statistical decisions in experimental and phenomenological analyses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。