提出实用导向的多分类校准评估框架,让模型预测更贴合实际决策需求。
Scalable Utility-Aware Multiclass Calibration
- 基于用户实用函数定义校准误差,统一多种现有评估方法。
- 可构建更鲁棒的顶部类别与类别级校准指标,适用于复杂决策场景。
- 适合关注模型可信度与实际应用效果的研究者和工程师。
确保分类器预测结果与实际观测频率一致是其具备可信性的基本要求。现有评估多分类校准的方法往往聚焦于特定预测特征(如最高类别置信度、类别级校准),或采用计算复杂的变分形式。本文研究可扩展的多分类校准评估方法,提出实用校准(utility calibration)框架:通过一个封装终端用户目标或决策标准的实用函数,衡量校准误差。该框架能统一并重新诠释多种现有校准指标,尤其支持更鲁棒的顶部类别与类别级校准指标,并进一步拓展至更丰富的下游实用场景评估。
原文摘要 · Abstract (English)
Ensuring that classifiers are well-calibrated, i.e., their predictions align with observed frequencies, is a minimal and fundamental requirement for classifiers to be viewed as trustworthy. Existing methods for assessing multiclass calibration often focus on specific aspects associated with prediction (e.g., top-class confidence, class-wise calibration) or utilize computationally challenging variational formulations. In this work, we study scalable \emph{evaluation} of multiclass calibration. To this end, we propose utility calibration, a general framework that measures the calibration error relative to a specific utility function that encapsulates the goals or decision criteria relevant to the end user. We demonstrate how this framework can unify and re-interpret several existing calibration metrics, particularly allowing for more robust versions of the top-class and class-wise calibration metrics, and, going beyond such binarized approaches, toward assessing calibration for richer classes of downstream utilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。