arXiv:2509.02279cs.LGcs.GT2025-09被引 4

用不可区分性重新定义校准,让概率预测更可信

Calibration through the Lens of Indistinguishability

  • 将校准视为预测世界与真实世界不可区分
  • 不同校准度量反映两类世界可被区分的程度
  • 适合关注预测可靠性与决策安全的研究者

校准是预测领域的一个经典概念,旨在回答:预测概率应如何解读?在仅能观测离散结果的世界中,如何评估一个对可能结果给出连续概率的预测器?由于机器学习中概率预测的普遍性,校准研究近年备受关注。本文综述了校准定义与误差测量的基础性问题,以及这些度量对依赖预测进行决策的下游使用者的意义。核心观点是:校准即不可区分性——预测所假设的世界与真实世界(由自然或贝叶斯最优预测器决定)无法被特定类别的区分者或统计方法分辨。各类校准度量正是量化这种可区分性的指标。

原文摘要 · Abstract (English)

Calibration is a classical notion from the forecasting literature which aims to address the question: how should predicted probabilities be interpreted? In a world where we only get to observe (discrete) outcomes, how should we evaluate a predictor that hypothesizes (continuous) probabilities over possible outcomes? The study of calibration has seen a surge of recent interest, given the ubiquity of probabilistic predictions in machine learning. This survey describes recent work on the foundational questions of how to define and measure calibration error, and what these measures mean for downstream decision makers who wish to use the predictions to make decisions. A unifying viewpoint that emerges is that of calibration as a form of indistinguishability, between the world hypothesized by the predictor and the real world (governed by nature or the Bayes optimal predictor). In this view, various calibration measures quantify the extent to which the two worlds can be told apart by certain classes of distinguishers or statistical measures.

校准概率预测不可区分性决策支持

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。