AI海平面预报集合在联合分布上存在偏差,虽单站准确但空间结构失真。
Do AI Forecast Ensembles Sample the Correct Conditional Distribution?

- 用扩散模型生成8个沿海站点的联合概率预报
- 单站预报有技巧,但空间相关性弱于气候参考值
- 该问题在能量评分下隐蔽,但变差分评分可检测
集合预报旨在采样结果的条件分布,但人工智能预报集合是否在联合意义上正确采样仍缺乏检验。本文针对美国东海岸8个验潮站,基于再分析数据训练扩散模型进行亚季节性沿海海平面概率预报。结果显示,边际预报质量与联合结构质量脱钩:所有站点和预报时效均具正技巧,但联合空间结构劣于气候随机抽样。通过置换分解发现,此失败对能量评分不敏感,但可被变差分评分检测。在0.7至170年等效训练量的Lorenz-96实验中,该差距持续存在,且线性基线亦复现此现象,表明学习到的分布存在结构性缺陷。动力集合未出现此问题,而确定性模拟器却重现了失败,说明该问题特异于学习型模拟器而非集合预报本身。
原文摘要 · Abstract (English)
Ensemble forecasting aims to sample the conditional distribution of outcomes; whether AI forecast ensembles do this correctly in a joint sense remains largely untested. We train a diffusion model for probabilistic subseasonal coastal sea level forecasts at eight US East Coast tide gauge stations, with sea level derived from reanalysis, and find that marginal and joint forecast quality decouple: positive skill at every station and lead time marginally, while joint spatial structure is worse than climatological draws. A shuffle-based permutation decomposition reveals this failure is invisible to the energy score but detected by the variogram score. Lorenz-96 experiments across 0.7-170 equivalent years show the gap persists regardless of training volume and is reproduced by a linear baseline, indicating structural inadequacy of the learned distribution. A dynamical ensemble does not replicate the failure while a deterministic emulator does, suggesting it is specific to learned emulators rather than ensemble forecasting generally.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。