提出可审计的多校准误差度量,解决模型可信度评估难题。
Auditability and the Landscape of Distance to Multicalibration
- 设计两种新度量,确保误差值反映修正模型所需改动量
- 证明现有方法在可审计性或修正意义上有缺陷
- 适用于需要高可信度公平性验证的AI系统开发者
校准是保证预测模型不确定性估计可信性的关键属性。多校准是对标准校准的强化,要求模型在可能重叠的多个子群体上均保持校准。随着多校准在实践中日益流行,一个核心问题是:如何衡量一个预测器的多校准程度?Błasiok等人(2023)通过距离校准框架(dCE)分析了标准校准度量之间的关系及与真实情况的关联。本文基于dCE框架,研究预测器的多校准距离的可审计性。我们考察了dCE向多子群体的两种自然推广:最坏组距离(wdMC)和多校准距离(dMC),并指出二者均不满足两个关键性质:1)度量应反映使模型完全多校准所需修改量;2)应在信息论意义上可审计。类似障碍也出现在一般多群体公平性度量的可审计性中。为此,我们提出了两种等价的多校准度量:一是dMC的连续化变体;二是基于交集公平性原则的距离到交集多校准。同时,我们揭示了多校准距离的损失景观以及完全多校准预测器集合的几何结构。这些发现对开发更强的多校准算法及多群体审计具有潜在影响。
原文摘要 · Abstract (English)
Calibration is a critical property for establishing the trustworthiness of predictors that provide uncertainty estimates. Multicalibration is a strengthening of calibration which requires that predictors be calibrated on a potentially overlapping collection of subsets of the domain. As multicalibration grows in popularity with practitioners, an essential question is: how do we measure how multicalibrated a predictor is? Błasiok et al. (2023) considered this question for standard calibration by introducing the distance to calibration framework (dCE) to understand how calibration metrics relate to each other and the ground truth. Building on the dCE framework, we consider the auditability of the distance to multicalibration of a predictor $f$. We begin by considering two natural generalizations of dCE to multiple subgroups: worst group dCE (wdMC), and distance to multicalibration (dMC). We argue that there are two essential properties of any multicalibration error metric: 1) the metric should capture how much $f$ would need to be modified in order to be perfectly multicalibrated; and 2) the metric should be auditable in an information theoretic sense. We show that wdMC and dMC each fail to satisfy one of these two properties, and that similar barriers arise when considering the auditability of general distance to multigroup fairness notions. We then propose two (equivalent) multicalibration metrics which do satisfy these requirements: 1) a continuized variant of dMC; and 2) a distance to intersection multicalibration, which leans on intersectional fairness desiderata. Along the way, we shed light on the loss-landscape of distance to multicalibration and the geometry of the set of perfectly multicalibrated predictors. Our findings may have implications for the development of stronger multicalibration algorithms as well as multigroup auditing more generally.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。