提出新校准度量方法,区分模型过自信与欠自信。
An Entropic Metric for Measuring Calibration of Machine Learning Models
- 基于目标跟踪理论构建熵校准差(ECD)度量
- 在真实与模拟数据上优于传统校准误差指标
- 适合关注模型可靠性与安全性的研究者
理解机器学习模型对输入样本分类时的置信度至关重要,但这一问题仍缺乏深入研究。本文提出一种新的校准度量——熵校准差(Entropic Calibration Difference, ECD),源自状态估计领域中的目标跟踪(TT)研究。该方法适用于二分类模型,能够有效区分模型的过自信与欠自信。在目标跟踪文献中,这两种偏差具有不同重要性,而传统校准度量常将其混淆。本研究强调:欠自信模型虽保守且统计效率较低,但更安全;而过自信模型存在风险。我们通过真实数据与模拟数据验证了ECD的有效性,并与期望校准误差(ECE)及其符号版本(ESCE)进行对比,结果表明ECD在识别校准偏差方面更具分辨力。
原文摘要 · Abstract (English)
Understanding the confidence with which a machine learning model classifies an input datum is an important, and perhaps under-investigated, concept. In this paper, we propose a new calibration metric, the Entropic Calibration Difference (ECD). Based on existing research in the field of state estimation, specifically target tracking (TT), we show how ECD may be applied to binary classification machine learning models. We describe the relative importance of under- and over-confidence and how they are not conflated in the TT literature. Indeed, our metric distinguishes under- from over-confidence. We consider this important given that algorithms that are under-confident are likely to be 'safer' than algorithms that are over-confident, albeit at the expense of also being over-cautious and so statistically inefficient. We demonstrate how this new metric performs on real and simulated data and compare with other metrics for machine learning model probability calibration, including the Expected Calibration Error (ECE) and its signed counterpart, the Expected Signed Calibration Error (ESCE).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。