arXiv:2410.23046cs.LG2024-10中稿 · AISTATS2025

提出无需真实不确定度即可评估模型可信度的新指标

Legitimate ground-truth-free metrics for deep uncertainty classification scoring

  • 基于测试数据设计可直接计算的不确定度评估指标
  • 证明这些指标与可解释的可信度排序存在理论关联
  • 适合关注模型安全性的深度学习实践者使用

尽管对更安全机器学习的需求日益增长,不确定性量化(UQ)方法在实际应用中的使用仍受限。这一限制主要源于缺乏不确定性真值进行验证。在分类任务中,当仅有常规测试数据时,多位研究者提出了仅依赖测试样本即可计算的评估指标来衡量量化不确定性质量。本文深入分析此类指标,证明其理论上表现良好,并与一种易于解释的不确定性真值相关联,该真值可直观反映模型预测的可信度排序。基于这些新发现,结合这些指标在常规监督学习范式下的适用性,本文认为其贡献将有助于推动深度学习中不确定性量化的更广泛应用。

原文摘要 · Abstract (English)

Despite the increasing demand for safer machine learning practices, the use of Uncertainty Quantification (UQ) methods in production remains limited. This limitation is exacerbated by the challenge of validating UQ methods in absence of UQ ground truth. In classification tasks, when only a usual set of test data is at hand, several authors suggested different metrics that can be computed from such test points while assessing the quality of quantified uncertainties. This paper investigates such metrics and proves that they are theoretically well-behaved and actually tied to some uncertainty ground truth which is easily interpretable in terms of model prediction trustworthiness ranking. Equipped with those new results, and given the applicability of those metrics in the usual supervised paradigm, we argue that our contributions will help promoting a broader use of UQ in deep learning.

不确定性量化可信度评估深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。