arXiv:2507.12094cs.LGcs.GT2025-07

提出新指标比较预测模型在决策中的实际价值,超越传统校准方法。

Is This Predictor More Informative than Another? A Decision-Theoretical Comparison

  • 定义预测模型间的信息差距,衡量其在各类决策任务中的最大收益优势
  • 实验显示该指标比传统评估方法更贴近真实决策效果
  • 适合关注预测模型落地价值的开发者与决策者使用

在诸多现实应用中,模型提供方为下游决策者生成概率预测,后者依据不同收益目标做出决策。提供方可能拥有多个可能存在偏差的预测模型,需选择最能提升决策效用的模型。核心挑战在于:当两个模型均未校准时,且决策任务因用户和场景而异,如何有效比较?本文首次提出‘信息差距’概念,即任一预测模型在所有决策任务中相对于另一模型的最大归一化收益优势。该框架严格推广了现有方法:既涵盖将有偏模型与校准版本对比的U-校准与校准决策损失,也包含双模型完全校准时的Blackwell信息性作为特例。第二项贡献是信息差距的对偶刻画,由此导出一种可视为预测分布间地球移动距离松弛版的信息度量。该度量满足完备性与合理性,且可在仅访问预测结果的设定下高效样本估计。通过基于大语言模型的预测器在真实任务上的实验验证,信息差距相比传统指标更具决策相关性,并为经验性校准后处理对决策效用的影响提供了严谨评估视角。

原文摘要 · Abstract (English)

In many real-world applications, a model provider provides probabilistic forecasts to downstream decision-makers who use them to make decisions under diverse payoff objectives. The provider may have access to multiple predictive models, each potentially miscalibrated, and must choose which model to deploy in order to maximize the usefulness of predictions for downstream decisions. A central challenge arises: how can the provider meaningfully compare two predictors when neither is guaranteed to be well-calibrated, and when the relevant decision tasks may differ across users and contexts? To answer this, our first contribution introduces the notion of the informativeness gap between any two predictors, defined as the maximum normalized payoff advantage one predictor offers over the other across all decision-making tasks. Our framework strictly generalizes several existing notions: it subsumes U-Calibration and Calibration Decision Loss, which compare a miscalibrated predictor to its calibrated counterpart, and it recovers Blackwell informativeness as a special case when both predictors are perfectly calibrated. Our second contribution is a dual characterization of the informativeness gap, which gives rise to a natural informativeness measure that can be viewed as a relaxed variant of the earth mover's distance between two prediction distributions. We show that this measure satisfies natural desiderata: it is complete and sound, and it can be estimated sample-efficiently in the prediction-only access setting. We complement our theory with experiments on LLM-based forecasters in real-world prediction tasks, showing that the informativeness gap offers a more decision-relevant alternative to traditional metrics, and provides a principled lens for evaluating how ad hoc calibration post-processing affects downstream decision usefulness.

预测评估决策优化信息差距

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。