提出可无需模型内部信息的多分类概率校准评估与修正方法
Multiclass Calibration Assessment and Recalibration of Probability Predictions via the Linear Log Odds Calibration Function
- 基于线性对数几率函数构建校准检验与修正框架
- 在图像、肥胖分析等三类真实场景中验证效果显著
- 结果易懂且适用于各类分类模型,尤其适合实际应用
机器生成的概率预测在现代分类任务(如图像分类)中至关重要。模型校准良好意味着其预测概率与实际事件频率一致。尽管需要多类别校准方法,现有方法存在三方面局限:(i) 仅能比较多个模型间的校准性能,无法直接评估单个模型;(ii) 需要访问模型内部信息(如神经网络层的logit输出);(iii) 输出难以被人类分析师理解。为克服上述问题,我们提出多类别线性对数几率(MCLLO)校准方法:(i) 引入似然比假设检验以评估校准性;(ii) 不需模型内部访问,适用于广泛分类任务;(iii) 结果易于解释。通过模拟和三个真实案例研究(卷积神经网络图像分类、随机森林肥胖分析、回归建模生态学)验证了MCLLO的有效性。与四种对比校准方法结合使用,分别采用本方法的假设检验与现有的期望校准误差(ECE)指标,证明该方法单独或协同使用均表现优异。
原文摘要 · Abstract (English)
Machine-generated probability predictions are essential in modern classification tasks such as image classification. A model is well calibrated when its predicted probabilities correspond to observed event frequencies. Despite the need for multicategory recalibration methods, existing methods are limited to (i) comparing calibration between two or more models rather than directly assessing the calibration of a single model, (ii) requiring under-the-hood model access, e.g., accessing logit-scale predictions within the layers of a neural network, and (iii) providing output which is difficult for human analysts to understand. To overcome (i)-(iii), we propose Multicategory Linear Log Odds (MCLLO) recalibration, which (i) includes a likelihood ratio hypothesis test to assess calibration, (ii) does not require under-the-hood access to models and is thus applicable on a wide range of classification problems, and (iii) can be easily interpreted. We demonstrate the effectiveness of the MCLLO method through simulations and three real-world case studies involving image classification via convolutional neural network, obesity analysis via random forest, and ecology via regression modeling. We compare MCLLO to four comparator recalibration techniques utilizing both our hypothesis test and the existing calibration metric Expected Calibration Error to show that our method works well alone and in concert with other methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。