提出新方法优化分类模型校准误差估计,提升模型可信度。
Optimizing Estimators of Squared Calibration Errors in Classification
- 将校准误差估计转为独立同分布回归问题,可量化评估不同方法。
- 在标准图像分类任务上验证,优化后方法显著降低校准误差。
- 适合关注模型可靠性与决策可信性的研究人员使用。
本文提出一种基于均方误差的风险函数,可在实际场景中比较和优化分类模型的平方校准误差估计器。提升分类器校准性能对增强机器学习模型的可信度和可解释性至关重要,尤其在敏感决策场景中。尽管已有多种校准(误差)估计器,但缺乏选择合适估计器及调参的指导。通过利用平方校准误差的双线性结构,我们将校准估计重构为独立同分布输入对的回归问题。该重构使我们能够即使在最具挑战性的“标准校准”准则下,也量化不同估计器的性能。我们的方法主张在评估数据集上采用训练-验证-测试流程来估计校准误差。通过优化现有估计器并对比基于核岭回归的新估计器,我们在标准图像分类任务上验证了该流程的有效性。
原文摘要 · Abstract (English)
In this work, we propose a mean-squared error-based risk that enables the comparison and optimization of estimators of squared calibration errors in practical settings. Improving the calibration of classifiers is crucial for enhancing the trustworthiness and interpretability of machine learning models, especially in sensitive decision-making scenarios. Although various calibration (error) estimators exist in the current literature, there is a lack of guidance on selecting the appropriate estimator and tuning its hyperparameters. By leveraging the bilinear structure of squared calibration errors, we reformulate calibration estimation as a regression problem with independent and identically distributed (i.i.d.) input pairs. This reformulation allows us to quantify the performance of different estimators even for the most challenging calibration criterion, known as canonical calibration. Our approach advocates for a training-validation-testing pipeline when estimating a calibration error on an evaluation dataset. We demonstrate the effectiveness of our pipeline by optimizing existing calibration estimators and comparing them with novel kernel ridge regression-based estimators on standard image classification tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。