arXiv:2506.17968cs.LGcs.AI2025-06TPAMI被引 3

提出h-calibration框架,用概率化目标实现更可靠的分类校准。

h-calibration: Rethinking Classifier Recalibration with Probabilistic Error-Bounded Objective

  • 基于概率误差有界的目标构建校准新框架
  • 在多个基准上达到当前最优校准性能
  • 适合需要高可靠性概率输出的模型部署场景

深度神经网络在众多任务中表现卓越,但常存在概率输出失准问题。为此,学界提出多种后处理校准方法,在不损害预训练模型分类性能的前提下提升概率可靠性。本文系统梳理并归类现有方法为三类:直观设计、分箱法和理想校准公式法。通过理论与实践分析,揭示了前人方法存在的十项共性缺陷。为此,我们提出h-calibration——一种基于概率误差有界目标的校准学习框架,理论上构建了标准校准的等价可微学习形式。在此基础上设计了一种简单高效的后处理校准算法。实验表明,该方法不仅克服全部十项缺陷,且显著优于传统方法。进一步的理论与实验分析证实,所提学习目标相较于传统合理评分规则具有更优性质。研究验证了该框架在标准后处理校准基准上的有效性,实现了最先进性能,为相关领域可靠置信度学习提供了重要参考。

原文摘要 · Abstract (English)

Deep neural networks have demonstrated remarkable performance across numerous learning tasks but often suffer from miscalibration, resulting in unreliable probability outputs. This has inspired many recent works on mitigating miscalibration, particularly through post-hoc recalibration methods that aim to obtain calibrated probabilities without sacrificing the classification performance of pre-trained models. In this study, we summarize and categorize previous works into three general strategies: intuitively designed methods, binning-based methods, and methods based on formulations of ideal calibration. Through theoretical and practical analysis, we highlight ten common limitations in previous approaches. To address these limitations, we propose a probabilistic learning framework for calibration called h-calibration, which theoretically constructs an equivalent learning formulation for canonical calibration with boundedness. On this basis, we design a simple yet effective post-hoc calibration algorithm. Our method not only overcomes the ten identified limitations but also achieves markedly better performance than traditional methods, as validated by extensive experiments. We further analyze, both theoretically and experimentally, the relationship and advantages of our learning objective compared to traditional proper scoring rule. In summary, our probabilistic framework derives an approximately equivalent differentiable objective for learning error-bounded calibrated probabilities, elucidating the correspondence and convergence properties of computational statistics with respect to theoretical bounds in canonical calibration. The theoretical effectiveness is verified on standard post-hoc calibration benchmarks by achieving state-of-the-art performance. This research offers valuable reference for learning reliable likelihood in related fields.

模型校准概率推理深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。