从决策者视角评估天气预报,发现传统方法可能选错最优模型。
Evaluating Weather Forecasts from a Decision Maker's Perspective
- 用决策校准框架评估预报在实际决策中的表现
- 机器学习模型在部分任务中优于传统数值模型
- 同一模型在不同决策任务中排名可能反转
标准天气预报评估侧重于预报员视角和统计对比,但实际应用中预报用于支持决策,因此从决策者视角出发更具意义。本文提出决策校准框架,以评估预报在决策层面的表现。我们比较了机器学习与经典数值天气预报模型在多种依赖天气的决策任务中的表现。结果表明,预报层面的性能差异不能可靠地转化为决策层面的优势:某些差异仅在决策层面显现,且不同任务中模型排名会发生变化。研究证实,常规评估无法有效选出特定决策任务的最优预报模型。
原文摘要 · Abstract (English)
Standard weather forecast evaluations focus on the forecaster's perspective and on a statistical assessment comparing forecasts and observations. In practice, however, forecasts are used to make decisions, so it seems natural to take the decision-maker's perspective and quantify the value of a forecast by its ability to improve decision-making. Decision calibration provides a novel framework for evaluating forecast performance at the decision level rather than the forecast level. We evaluate decision calibration to compare Machine Learning and classical numerical weather prediction models on various weather-dependent decision tasks. We find that model performance at the forecast level does not reliably translate to performance in downstream decision-making: some performance differences only become apparent at the decision level, and model rankings can change among different decision tasks. Our results confirm that typical forecast evaluations are insufficient for selecting the optimal forecast model for a specific decision task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。