扩展概率评分与校准理论,解决模糊概率预测的评估难题
Scoring Rules and Calibration for Imprecise Probabilities
- 将精确概率的评分规则和校准概念推广至概率集合场景
- 揭示模糊预测中评分与校准目标可能不一致的深层原因
- 为分布鲁棒性学习中的损失函数选择提供理论警示
当说明天气雨的概率在20%到30%之间时,如何评估这类模糊概率预测?精确概率预测的评估理论已成熟,基于合理评分规则和校准性。但针对概率集合(imprecise probabilities)的情形,相关理论仍不完善。本文将合理评分规则与校准性概念推广至模糊情形,将其建立在数据模型与决策问题的上下文中,使不确定性具有明确语境。研究揭示了其与(群体)分布鲁棒性范式的紧密联系,并提出:合理评分规则与校准在精确情形下目标一致,但在模糊情形下可能分离。决策论熵在两者中起关键作用。最后,通过机器学习实践验证理论,揭示分布鲁棒性中损失函数选择的潜在陷阱。
原文摘要 · Abstract (English)
What does it mean to say that, for example, the probability for rain tomorrow is between 20% and 30%? The theory for the evaluation of precise probabilistic forecasts is well-developed and is grounded in the key concepts of proper scoring rules and calibration. For the case of imprecise probabilistic forecasts (sets of probabilities), such theory is still lacking. In this work, we therefore generalize proper scoring rules and calibration to the imprecise case. We develop these concepts as relative to data models and decision problems. As a consequence, the imprecision is embedded in a clear context. We establish a close link to the paradigm of (group) distributional robustness and in doing so provide new insights for it. We argue that proper scoring rules and calibration serve two distinct goals, which are aligned in the precise case, but intriguingly are not necessarily aligned in the imprecise case. The concept of decision-theoretic entropy plays a key role for both goals. Finally, we demonstrate the theoretical insights in machine learning practice, in particular we illustrate subtle pitfalls relating to the choice of loss function in distributional robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。