提出新方法评估模型校准度,能更准确比较不同模型优劣。
All Models Are Miscalibrated, But Some Less So: Comparing Calibration with Conditional Mean Operators
- 基于条件均值算子的希尔伯特-施密特范数,直接衡量条件分布差异
- 在合成与真实数据上,对模型校准度排序更一致且抗分布偏移
- 适合高风险场景中需精确评估预测置信度的研究者使用
在高风险应用中,概率预测模型的校准度至关重要。然而,现有校准误差估计器常无法有效区分哪个模型更优。本文提出条件核校准误差(CKCE),基于希尔伯特-施密特范数衡量条件均值算子之差。通过直接利用强校准的定义——即条件分布间的距离,并以再生核希尔伯特空间中的嵌入表示,CKCE 对预测模型的边缘分布不敏感,相较以往指标更适用于相对比较。实验结果表明,无论在合成数据还是真实数据上,CKCE 均能提供更一致的模型校准度排名,且对分布偏移更具鲁棒性。
原文摘要 · Abstract (English)
When working in a high-risk setting, having well calibrated probabilistic predictive models is a crucial requirement. However, estimators for calibration error are not always able to correctly distinguish which model is better calibrated. We propose the \emph{conditional kernel calibration error} (CKCE) which is based on the Hilbert-Schmidt norm of the difference between conditional mean operators. By working directly with the definition of strong calibration as the distance between conditional distributions, which we represent by their embeddings in reproducing kernel Hilbert spaces, the CKCE is less sensitive to the marginal distribution of predictive models. This makes it more effective for relative comparisons than previously proposed calibration metrics. Our experiments, using both synthetic and real data, show that CKCE provides a more consistent ranking of models by their calibration error and is more robust against distribution shift.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。