arXiv:2508.13100cs.LGcs.DS2025-08被引 6

提出首个批量设置下完全诚实的校准度量,可准确评估预测概率可靠性。

A Perfectly Truthful Calibration Measure

  • 设计基于双区间误差的简单校准度量ATB,确保真实概率时最小化
  • 首次实现线性时间校准检验,比此前方法更快捷高效
  • 提供通用构建诚实度量的方法,适用于多种校准场景

校准要求预测条件无偏,从而可靠地解释为概率。校准度量用于量化预测器与完美校准之间的距离。根据Haghtalab等人(2024)的定义,若预测器输出真实概率时期望值最小,则该度量为诚实的。然而,在随机样本上评估时,现有所有校准度量都会激励预测器说谎以显得更校准。尽管已有研究在顺序预测中构造近似诚实的度量,但在更基础的批量设置下,仍不存在完全诚实的校准度量。本文提出一种简单、严格且完全诚实的批量校准度量——平均双区间校准误差(ATB),其与现有两个度量(smooth calibration error, smCal;lower distance to calibration, distCal)呈二次关系。ATB定义简洁,计算高效,首次实现线性时间校准测试,优于Hu等(2024)的结果。此外,我们提出基于独立随机变量方差可加性的通用构造方法,不仅证明了ATB的诚实性,还可推广至其他度量如分位数分箱l_2-ECE。

原文摘要 · Abstract (English)

Calibration requires that predictions are conditionally unbiased and, therefore, reliably interpretable as probabilities. A calibration measure quantifies how far a predictor is from perfect calibration. As introduced by Haghtalab et al. (2024), a calibration measure is truthful if it is minimized in expectation when a predictor outputs the ground-truth probabilities. Predicting the true probabilities guarantees perfect calibration, but in reality, when calibration is evaluated on a random sample, all known calibration measures incentivize predictors to lie in order to appear more calibrated. Such lack of truthfulness motivated Haghtalab et al. (2024) and Qiao and Zhao (2025) to construct approximately truthful calibration measures in the sequential prediction setting, but no perfectly truthful calibration measure was known to exist even in the more basic batch setting. We design a simple, perfectly and strictly truthful, sound and complete calibration measure in the batch setting: averaged two-bin calibration error (ATB). ATB is quadratically related to two existing calibration measures: the smooth calibration error smCal and the lower distance to calibration distCal. The simplicity in our definition of ATB makes it efficient and straightforward to compute, allowing us to give the first linear-time calibration testing algorithm, improving a result of Hu et al. (2024). We also introduce a general recipe for constructing truthful measures based on the variance additivity of independent random variables, which proves the truthfulness of ATB as a special case and allows us to construct other truthful calibration measures such as quantile-binned l_2-ECE.

校准度量概率预测诚实性算法效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。