arXiv:2504.18278cs.LGstat.ML2025-04综述被引 6

系统梳理82种概率校准指标,帮模型更真实地表达预测信心。

A comprehensive review of classifier probability calibration metrics

  • 按点、分箱、核函数等四类方法分类整理校准指标
  • 涵盖82个核心指标,提供公式支持可复现
  • 适合关注模型可信度与安全应用的研究者

人工智能和机器学习模型输出的概率或置信度常与实际准确率不符,例如模型声称80%确定时,是否真的正确80%?概率校准指标用于衡量置信度与真实准确率之间的偏差,独立评估模型的校准性能,补充传统准确率指标。理解校准在多系统融合、安全或商业关键场景中至关重要,也有助于建立用户对模型的信任。本文全面综述了分类器与目标检测模型的概率校准指标,依据多种分类标准进行组织,揭示其内在关联。共识别出82个主要指标,可分为四类分类器家族(点式、分箱式、核函数/曲线式、累积式)及一类目标检测家族。对每个指标均提供可用公式,便于后续研究实现与比较。

原文摘要 · Abstract (English)

Probabilities or confidence values produced by artificial intelligence (AI) and machine learning (ML) models often do not reflect their true accuracy, with some models being under or over confident in their predictions. For example, if a model is 80% sure of an outcome, is it correct 80% of the time? Probability calibration metrics measure the discrepancy between confidence and accuracy, providing an independent assessment of model calibration performance that complements traditional accuracy metrics. Understanding calibration is important when the outputs of multiple systems are combined, for assurance in safety or business-critical contexts, and for building user trust in models. This paper provides a comprehensive review of probability calibration metrics for classifier and object detection models, organising them according to a number of different categorisations to highlight their relationships. We identify 82 major metrics, which can be grouped into four classifier families (point-based, bin-based, kernel or curve-based, and cumulative) and an object detection family. For each metric, we provide equations where available, facilitating implementation and comparison by future researchers.

概率校准模型可信度评估指标综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。