arXiv:2602.12975cs.LGcs.AI2026-02被引 1

提出新指标VCE,可评估模型对概率分布变异的校准能力。

Extending confidence calibration to generalised measures of variation

  • 将经典置信度校准扩展至任意变异度量,支持全概率分布分析。
  • 在设计为完美校准的合成数据上,VCE随样本数增加趋近零。
  • 相比已有熵基指标UCE,VCE具备更优收敛性,适合严谨校准评估。

我们提出一种新的变异校准误差(Variation Calibration Error, VCE)指标,用于评估机器学习分类器的校准性能。该指标可视为广为人知的期望校准误差(Expected Calibration Error, ECE)的拓展,后者仅衡量最大概率或置信度的校准情况。其他能考虑完整概率分布的变异度量,如香农熵,具有优势。我们展示了如何将ECE方法从置信度校准推广到任意变异度量的校准评估。通过合成预测的数值示例,这些预测按设计是完美校准的,结果表明:随着数据样本数量增加,VCE趋近于零;而文献中已有的基于熵的校准指标(即不确定性校准误差,UCE)则不具备此性质。

原文摘要 · Abstract (English)

We propose the Variation Calibration Error (VCE) metric for assessing the calibration of machine learning classifiers. The metric can be viewed as an extension of the well-known Expected Calibration Error (ECE) which assesses the calibration of the maximum probability or confidence. Other ways of measuring the variation of a probability distribution exist which have the advantage of taking into account the full probability distribution, for example the Shannon entropy. We show how the ECE approach can be extended from assessing confidence calibration to assessing the calibration of any metric of variation. We present numerical examples upon synthetic predictions which are perfectly calibrated by design, demonstrating that, in this scenario, the VCE has the desired property of approaching zero as the number of data samples increases, in contrast to another entropy-based calibration metric (the UCE) which has been proposed in the literature.

模型校准变异度量评估指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。