arXiv:2504.15582cs.LGcs.DS2025-04被引 8

改进预测校准方法,让机器决策更可靠。

Smooth Calibration and Decision Making

  • 通过加噪实现差分隐私,优化决策校准误差
  • 在线预测器经后处理后ECE与CDL降至O(√ε)
  • 适合关注决策可靠性而非单纯模型精度的场景

校准要求预测结果与其贝叶斯后验一致。对于不区分微小扰动的机器学习预测器,校准误差在预测上是连续的,如平滑校准误差(Foster and Hart, 2018)和距离校准(Blasiok et al., 2023a)。相反,使用预测进行决策的决策者在概率空间中做的是非连续最优决策,其因校准错误产生的损失也是非连续的。因此,针对机器学习的低校准误差预测器可能在决策中表现出高校准误差,即对假设预测正确的决策者不可靠。本文研究:能否对一个与校准距离为ε的在线预测器进行后处理而不损失,从而实现低决策校准误差?我们证明,此类后处理可达到O(√ε)的期望校准误差(ECE)和校准决策损失(CDL),且该界渐近最优。后处理算法引入噪声以实现差分隐私。然而,从低校准距离预测器出发的后处理最优界仍不如直接优化ECE和CDL的现有在线校准算法。

原文摘要 · Abstract (English)

Calibration requires predictor outputs to be consistent with their Bayesian posteriors. For machine learning predictors that do not distinguish between small perturbations, calibration errors are continuous in predictions, e.g., smooth calibration error (Foster and Hart, 2018), Distance to Calibration (Blasiok et al., 2023a). On the contrary, decision-makers who use predictions make optimal decisions discontinuously in probabilistic space, experiencing loss from miscalibration discontinuously. Calibration errors for decision-making are thus discontinuous, e.g., Expected Calibration Error (Foster and Vohra, 1997), and Calibration Decision Loss (Hu and Wu, 2024). Thus, predictors with a low calibration error for machine learning may suffer a high calibration error for decision-making, i.e., they may not be trustworthy for decision-makers optimizing assuming their predictions are correct. It is natural to ask if post-processing a predictor with a low calibration error for machine learning is without loss to achieve a low calibration error for decision-making. In our paper, we show that post-processing an online predictor with $ε$ distance to calibration achieves $O(\sqrtε)$ ECE and CDL, which is asymptotically optimal. The post-processing algorithm adds noise to make predictions differentially private. The optimal bound from low distance to calibration predictors from post-processing is non-optimal compared with existing online calibration algorithms that directly optimize for ECE and CDL.

校准决策差分隐私

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。