arXiv:2606.10777cs.LG2026-06

提出新评估标准,检验模型对认知不确定性估计的可靠性

Can we trust our models? Epistemic calibration in second-order classification

论文配图:Can we trust our models? Epistemic calibration in second-order classification
图 1 · 摘自论文原文
  • 引入认知校准概念,衡量预测波动与真实值的匹配度
  • 实验证明多种方法在预测性能相似下,认知校准差异显著
  • 适合高风险场景中需可信不确定性的研究者使用

不确定性估计对机器学习在高风险场景中的部署至关重要。传统校准仅评估预测概率的可靠性,无法判断认知不确定性估计本身是否可信。这一局限在二阶分类模型中尤为突出。本文提出认知校准(epistemic calibration),一种衡量报告的认知不确定性是否真实反映模型预测围绕真实值的分布程度的原则性标准。我们证明认知校准严格强于传统校准,能捕捉标准指标无法察觉的失效模式。通过一个在认知校准假设下成立的不可能性定理,连接已有文献。为可操作化该概念,提出期望认知校准误差(EECE),并证明其是真认知校准误差(TECE)的一致估计器。跨多种不确定性量化方法的实验表明,认知校准是一个一致且有意义的标准,揭示了即使预测性能相近的方法间也存在显著差异。

原文摘要 · Abstract (English)

Uncertainty estimation is critical for deploying machine learning models in high-stakes settings. However, classical calibration only assesses the reliability of predicted probabilities and does not evaluate whether epistemic uncertainty estimates are themselves trustworthy. This limitation is particularly relevant for second-order classification models. We introduce epistemic calibration, a principled criterion that measures whether reported epistemic uncertainty faithfully reflects the dispersion of model predictions around the ground truth. We show that epistemic calibration is a strictly stronger notion than classical calibration and captures failure modes invisible to standard metrics. We relate this work to the existing literature through an impossibility theorem that holds under the epistemic calibration hypothesis. To operationalize this concept, we propose the Expected Epistemic Calibration Error (EECE), which we prove to be a consistent estimator of a True Epistemic Calibration Error (TECE). Experiments across a broad range of uncertainty quantification methods show that epistemic calibration is a coherent and meaningful criterion and reveal substantial differences across methods, despite similar predictive performance.

不确定性估计模型校准二阶分类可信AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。