提出新指标CSR,更好识别模型过自信风险
Beyond ECE: Calibrated Size Ratio, Risk Assessment, and Confidence-Weighted Metrics

- 用可解释的校准尺寸比CSR替代传统ECE
- 发现标准方法校准后仍可能产生高风险过自信
- 引入置信度加权指标,提升分类性能评估精度
置信度校准长期依赖期望校准误差(ECE),该指标对所有置信度水平下的校准偏差同等对待。我们发现ECE即使数值很小,也可能存在任意大的过自信风险,因此提出可解释的校准尺寸比(CSR),其在完美校准时等于1,并由此推导出过自信风险概率 $P_{\mathrm{risk}}$,量化过自信的统计证据。我们进一步指出,过自信风险评估需与判别价值衡量结合:置信度是否能有效区分正确与错误预测。我们证明置信度加权准确率(cwA)是自然的补充,并推广至所有标准分类指标。特别地,我们证明置信度加权AUC(cwAUC)捕捉了校准信息,而经典AUC无法做到。我们在多个合成置信分布和十五个真实数据集上验证,发现CSR能有效区分有风险与无风险的置信分配,且标准后处理校准方法仍可能产生高风险置信输出。
原文摘要 · Abstract (English)
Confidence calibration has been dominated by the Expected Calibration Error (ECE), a linear metric that counts calibration offset equally regardless of the confidence level at which it occurs. We show that ECE can remain small even under arbitrarily large overconfidence risk, so we propose Calibrated Size Ratio (CSR) instead, an interpretable metric that equals 1 under perfect calibration, from which we derive the risk probability $P_{\mathrm{risk}}$ that quantifies the statistical evidence for overconfidence. We further argue that overconfidence risk assessment must be complemented by a measure of discriminative value: whether the assigned confidences actively distinguish correct from incorrect predictions. We show that confidence-weighted accuracy $\mathrm{cwA}$ is the natural such complement, and that confidence-weighting extends to all standard classification metrics. In particular, we prove that the confidence-weighted AUC (cwAUC) captures the information about calibration while the classical AUC cannot. We validate the proposed indicators on several synthetic confidence distributions under multiple controlled calibration profiles and find that CSR separates risky from non-risky assignments. We also test the metrics on fifteen real datasets, with and without post-hoc calibration, and find that standard methods can yield risky confidence profiles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。