提出新指标衡量多校准程度,避免传统方法的噪声问题。
Measuring multi-calibration
- 基于库皮尔统计量,按信噪比加权子群体贡献。
- 在基准数据集上验证,未考虑信噪比时指标显著变噪。
- 适合关注公平性与预测可靠性研究者使用。
理想的概率预测在所有子群体中均应满足预测概率等于实际观测期望值,称为完全多校准。现实中预测很少达到完全多校准,因此需要一个度量距离完全多校准程度的统计量。本文提出一种基于经典库皮尔统计量的新指标,用于衡量多校准水平,该方法避免了分箱或核密度估计带来的常见问题。新指标按子群体的信噪比对贡献进行加权;消融实验表明,若省略信噪比权重,指标将变得高度嘈杂。数值实验在多个基准数据集上展示了该度量的有效性。
原文摘要 · Abstract (English)
A suitable scalar metric can help measure multi-calibration, defined as follows. When the expected values of observed responses are equal to corresponding predicted probabilities, the probabilistic predictions are known as "perfectly calibrated." When the predicted probabilities are perfectly calibrated simultaneously across several subpopulations, the probabilistic predictions are known as "perfectly multi-calibrated." In practice, predicted probabilities are seldom perfectly multi-calibrated, so a statistic measuring the distance from perfect multi-calibration is informative. A recently proposed metric for calibration, based on the classical Kuiper statistic, is a natural basis for a new metric of multi-calibration and avoids well-known problems of metrics based on binning or kernel density estimation. The newly proposed metric weights the contributions of different subpopulations in proportion to their signal-to-noise ratios; data analyses' ablations demonstrate that the metric becomes noisy when omitting the signal-to-noise ratios from the metric. Numerical examples on benchmark data sets illustrate the new metric.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。