arXiv:2504.18395cs.LG2025-04被引 6

厘清三种校准定义及其关系,助你理解预测系统可信性

Three Types of Calibration with Properties and their Semantic and Formal Relationships

  • 从预测属性自实现与损失精准估计出发,提出两类校准原型
  • 二元结果下两种定义等价,高维结果可统一为分布校准
  • 揭示校准与后悔最小化、多校准的关系,适合研究可信AI者阅读

随着对‘可信性’和算法公平性的关注,预测系统的校准问题重获学界重视。传统校准定义为:当预报降雨概率为p的天数中,实际下雨比例也为p。然而近年涌现出大量不同含义的校准概念,彼此不可比、目的各异或相互蕴含。本文从两个动机出发构建校准:一是预测属性的自我实现(基于反射原则),二是决策者损失的精确估计(基于精算公平)。针对结果分布的属性(如均值、中位数)提出原型定义——Γ-校准,在特定条件下等价于某种交换后悔;另一定义源自Zhao等人[73]的决策校准,用于损失估计。在二元结果下,两者在合适属性选择下一致;高维结果时,可通过分布校准统一。最后讨论分组在两类校准中的作用,常用于实现多校准。本文提供校准概念的语义地图,以梳理纷繁定义。

原文摘要 · Abstract (English)

Fueled by discussions around "trustworthiness" and algorithmic fairness, calibration of predictive systems has regained scholars attention. The vanilla definition and understanding of calibration is, simply put, on all days on which the rain probability has been predicted to be p, the actual frequency of rain days was p. However, the increased attention has led to an immense variety of new notions of "calibration." Some of the notions are incomparable, serve different purposes, or imply each other. In this work, we provide two accounts which motivate calibration: self-realization of forecasted properties and precise estimation of incurred losses of the decision makers relying on forecasts. We substantiate the former via the reflection principle and the latter by actuarial fairness. For both accounts we formulate prototypical definitions via properties $Γ$ of outcome distributions, e.g., the mean or median. The prototypical definition for self-realization, which we call $Γ$-calibration, is equivalent to a certain type of swap regret under certain conditions. These implications are strongly connected to the omniprediction learning paradigm. The prototypical definition for precise loss estimation is a modification of decision calibration adopted from Zhao et al. [73]. For binary outcome sets both prototypical definitions coincide under appropriate choices of reference properties. For higher-dimensional outcome sets, both prototypical definitions can be subsumed by a natural extension of the binary definition, called distribution calibration with respect to a property. We conclude by commenting on the role of groupings in both accounts of calibration often used to obtain multicalibration. In sum, this work provides a semantic map of calibration in order to navigate a fragmented terrain of notions and definitions.

校准可信AI决策公平概率预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。