arXiv:2510.06388cs.LGcs.DS2025-10被引 2

提出可信赖的多分类校准误差,让预测值更真实可信。

Truthful Calibration Errors for Multi-Class Prediction

  • 设计基于真实分布的校准误差度量,避免虚假优化
  • 新方法在不同分箱下排名稳定,解决传统方法波动问题
  • 适合需要可靠概率输出的模型评估场景

校准预测能将数值解释为概率,因此校准误差被广泛用于评估、比较和调优概率预测器。近期,Haghtalab 等(2024)提出校准度量需满足“真实性”:即预测器最小化期望误差时应报告真实的条件标签分布。许多标准经验校准误差不具备真实性:预测器可通过扭曲概率来显得更校准。本文研究多分类预测中真实性在校准测量中的实际作用。首先,我们为标签分布的多维线性性质引入完全真实校准误差,推广了 Hartline 等(2025)的二分类真实校准误差,涵盖全类别校准和类别级校准。我们还提出一种真实修正的置信度校准方法。其次,刻画了这些真实误差的决策论意义:对于已校准预测器,真实校准误差保持 Blackwell 支配性——更信息丰富的校准预测器不会获得更大期望误差。第三,该决策论解释可说明并缓解广为人知的分箱校准误差排名不稳定性问题。实验表明,非真实置信度误差在分箱数量变化时会颠倒模型排名,而我们的真实误差在不同分箱选择下给出更稳定的排名。

原文摘要 · Abstract (English)

Calibrated predictions are useful because their numerical values can be interpreted as probabilities. Calibration errors are therefore widely used to evaluate, compare, and tune probabilistic predictors. Recently, Haghtalab et al. (2024) introduced an additional requirement for such measures: truthfulness. A calibration measure is truthful if a predictor minimizes its expected measured error by reporting the true conditional label distribution. Many standard empirical calibration errors are non-truthful: a predictor may appear better calibrated by distorting its probabilities rather than reporting them truthfully. We study the practical role of truthfulness for calibration measurement in multiclass prediction. First, we introduce perfectly truthful calibration errors for multidimensional linear properties of the label distribution, generalizing the truthful calibration error for binary predictions in Hartline et al. (2025). This framework includes full multiclass calibration and classwise calibration. We also identify a truthful correction for confidence calibration. Second, we characterize the decision-theoretic implications of these truthful errors. For calibrated predictors, truthful calibration errors preserve the Blackwell dominance: a more informative calibrated predictor receives no larger expected error. Third, we show that this decision-theoretic interpretation explains and mitigates the well-observed ranking robustness problem of binned calibration errors. Empirically, non-truthful confidence-based errors can reverse model rankings when the number of bins changes, while our truthful errors give more stable rankings across binning choices.

校准误差多分类概率预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。