arXiv:2605.17749cs.LGstat.ML2026-05被引 2

提出新校准度量SCDL,兼具可行动与可测试性。

Testable and Actionable Calibration for Full Swap Regret

论文配图:Testable and Actionable Calibration for Full Swap Regret
图 1 · 摘自论文原文
  • 引入软分箱校准损失(SCDL),直接关联预测误差与决策损失。
  • 理论证明其在小样本下可准确估计,误差接近最优。
  • 适合需高可信度预测的医疗、金融等关键决策场景。

AI生成的预测越来越多地用于关键任务决策,因此必须具备可信性。校准是衡量可信性的常用指标,要求预测值与真实发生频率一致,可视为真实概率。然而校准定义微妙,设计有效的校准误差度量是近年研究热点。核心挑战在于:既要可行动——能告诉决策者将预测当作真实概率时的效用损失(即交换后悔值);又要可测试——仅凭少量预测与实际结果即可评估校准误差。现有方法均无法同时满足这两点:要么弱化可行动性(仅界定了更弱的交换后悔值),要么牺牲可测试性(估计误差较大)。本文提出一种新校准度量——软分箱校准决策损失(SCDL),证明其在不削弱任一要求的前提下完全可行动且可测试,且估计误差近乎最优。此外,SCDL还满足连续性与一致性等理想性质。实验验证了其理论优势在实践中也带来更好的表现。

原文摘要 · Abstract (English)

AI generated predictions increasingly inform decision making in critical tasks, and therefore must be trustworthy. One widely used measure of trustworthiness is calibration, which requires that the predictions match the true frequencies and can be treated like real probabilities of a given outcome. However, defining calibration is subtle, and designing good measures of calibration error has been an active topic of recent research. The first goal is to find calibration measures that are actionable, meaning they can inform decision makers about their utility loss when predictions are treated as true probabilities, which is known as swap regret. The second goal is to find calibration measures that are testable, meaning that calibration error can be measured from a small sample of predictions and outcomes. Although these are very basic requirements, there is no existing calibration measure that fully satisfies both properties, and all existing measures relax actionability by bounding a weaker notion of swap regret, or relax testability by having suboptimal estimation error. We introduce a new calibration measure, Soft-Binned Calibration Decision Loss (SCDL), which we prove is fully actionable without weakening either requirement, and testable with nearly optimal error rate. In addition, SCDL satisfies other desired properties such as continuity and consistency. We also provide a set of experiments confirming that the theoretical advantages of SCDL compared to other measures lead to better performance in practice.

校准决策可测试性交换后悔

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。