量化分类器校准前后的决策风险,指导是否值得进一步优化。
Decision from Suboptimal Classifiers: Excess Risk Pre- and Post-Calibration
- 用校准曲线和分组损失估计器计算决策风险
- 发现校准可解决大部分风险,但部分场景需更深层优化
- 为模型后处理提供成本效益判断依据,适合工程落地
概率分类器在不确定性下的决策中至关重要。基于最大期望效用原则,最优决策规则可由后验类别概率和误分类代价推导得出。然而实践中仅能获得近似的真实后验概率。本文量化了在批量二分类决策中使用近似后验概率带来的额外风险(即后悔值)。我们给出了由校准误差引起的后悔值 $R^{ ext{CL}}$ 的解析表达式,以及对校准后分类器后悔值 $R^{ ext{GL}}$ 的紧致上下界。这些表达式揭示了仅校准即可缓解大部分风险的场景,也识别出受分组损失主导的场景,提示需超越校准的后训练策略。关键的是,$R^{ ext{CL}}$ 和 $R^{ ext{GL}}$ 均可通过校准曲线与近期提出的分组损失估计器在实际中估算。在 NLP 实验中,我们验证了这些指标能有效判断更高级后训练是否值得投入。最后,我们指出多校准方法可作为更昂贵微调的高效替代方案。
原文摘要 · Abstract (English)
Probabilistic classifiers are central for making informed decisions under uncertainty. Based on the maximum expected utility principle, optimal decision rules can be derived using the posterior class probabilities and misclassification costs. Yet, in practice only learned approximations of the oracle posterior probabilities are available. In this work, we quantify the excess risk (a.k.a. regret) incurred using approximate posterior probabilities in batch binary decision-making. We provide analytical expressions for miscalibration-induced regret ($R^{\mathrm{CL}}$), as well as tight and informative upper and lower bounds on the regret of calibrated classifiers ($R^{\mathrm{GL}}$). These expressions allow us to identify regimes where recalibration alone addresses most of the regret, and regimes where the regret is dominated by the grouping loss, which calls for post-training beyond recalibration. Crucially, both $R^{\mathrm{CL}}$ and $R^{\mathrm{GL}}$ can be estimated in practice using a calibration curve and a recent grouping loss estimator. On NLP experiments, we show that these quantities identify when the expected gain of more advanced post-training is worth the operational cost. Finally, we highlight the potential of multicalibration approaches as efficient alternatives to costlier fine-tuning approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。