用二项过程建模校准数据,提升分类置信度校准精度。
Combining Priors with Experience: Confidence Calibration Based on Binomial Process Modeling
- 将校准数据采样视为二项过程,联合先验与经验数据估计连续校准曲线。
- 所需样本量仅为直方图分箱法的 $3/B$,且校准曲线满足Lipschitz连续性。
- 提出新度量 $TCE_{bpm}$,可一致估计真实校准误差,适合模型评估与比较。
分类模型的置信度校准旨在估计预测类别的真实后验概率,对实际应用中的可靠决策至关重要。现有方法多基于统计技术从数据中估计校准曲线或拟合用户定义的函数,但常忽略校准曲线背后的先验分布。本文提出一种新方法,将校准曲线的先验分布与实证数据结合,通过建模校准数据采样过程为二项过程并最大化其似然函数,实现连续校准曲线估计。证明该方法在数据分布上具有Lipschitz连续性,且所需样本量仅为直方图分箱法的 $3/B$($B$ 为分箱数)。此外,设计新校准度量 $TCE_{bpm}$,利用估计校准曲线估计真实校准误差(TCE),并证明其为一致度量。还可通过预设真实校准曲线与置信度分布,生成真实校准数据集,作为基准评估现有度量与真实误差的差异。在真实与模拟数据上验证了方法与度量的有效性。
原文摘要 · Abstract (English)
Confidence calibration of classification models is a technique to estimate the true posterior probability of the predicted class, which is critical for ensuring reliable decision-making in practical applications. Existing confidence calibration methods mostly use statistical techniques to estimate the calibration curve from data or fit a user-defined calibration function, but often overlook fully mining and utilizing the prior distribution behind the calibration curve. However, a well-informed prior distribution can provide valuable insights beyond the empirical data under the limited data or low-density regions of confidence scores. To fill this gap, this paper proposes a new method that integrates the prior distribution behind the calibration curve with empirical data to estimate a continuous calibration curve, which is realized by modeling the sampling process of calibration data as a binomial process and maximizing the likelihood function of the binomial process. We prove that the calibration curve estimating method is Lipschitz continuous with respect to data distribution and requires a sample size of $3/B$ of that required for histogram binning, where $B$ represents the number of bins. Also, a new calibration metric ($TCE_{bpm}$), which leverages the estimated calibration curve to estimate the true calibration error (TCE), is designed. $TCE_{bpm}$ is proven to be a consistent calibration measure. Furthermore, realistic calibration datasets can be generated by the binomial process modeling from a preset true calibration curve and confidence score distribution, which can serve as a benchmark to measure and compare the discrepancy between existing calibration metrics and the true calibration error. The effectiveness of our calibration method and metric are verified in real-world and simulated data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。