arXiv:2609.08412cs.LGphysics.ao-ph2026-09

给已有天气预测模型加随机扰动,低成本生成不确定性预报。

Stochastically Perturbed Weights: Ensembles from Deterministic Machine-Learning Weather Models

论文配图:Stochastically Perturbed Weights: Ensembles from Deterministic Machine-Learning Weather Models
图 1 · 摘自论文原文
  • 推理时随机扰动网络权重,无需重新训练即可生成概率预报
  • 10天预报时,性能仅比专用概率模型低0.04~0.13的CRPSS分数
  • 适用于需要快速获取不确定性估计的气象业务场景

机器学习天气模型(MLWMs)在中长期全球预报上已达到或超越传统数值天气预报(NWP),且推理成本更低。多数部署的MLWM为确定性模型,无法提供自身预测的不确定性,而概率模型虽可生成校准的集合预报,但需额外训练。本文探讨如何从已有确定性模型中提取不确定性,而不必重新训练。受物理集合通过随机扰动参数化倾向启发,本文提出在推理时随机扰动网络原始权重(称为SPW)。在四个确定性骨干模型(Aurora、GraphCast、SFNO、AIFS)上进行三阶段消融实验,对比训练好的概率模型AIFS-ENS、FourCastNet 3、Atlas及欧洲中期天气预报中心(ECMWF)集合(IFS-ENS),覆盖112次初始时刻。在240小时(10天)预报时效下,SPW集合的连续排名概率技能得分(CRPSS)比最优概率模型低0.04至0.13,且无额外训练成本。不同模型的最佳噪声注入位置各异,说明该方法依赖架构特性,目前需调参而非即插即用。主要失败表现为全域均值系统性偏移,导致过度离散;限制噪声作用于粗尺度或扰动初值可部分修复此问题。

原文摘要 · Abstract (English)

Machine-learning weather models (MLWMs) now match or outperform operational numerical weather prediction (NWP) at global medium-range forecasting, at far lower inference cost. Many deployed MLWMs are deterministic, producing a single forecast with no estimate of its own uncertainty, whereas a growing family of trained-probabilistic models generate calibrated ensembles directly, at the price of a dedicated training run. We ask instead how much uncertainty can be extracted from a deterministic checkpoint that already exists, without retraining it. Where physical ensembles represent model uncertainty by stochastically perturbing parametrisation tendencies, we perturb the network's raw weight tensors at inference time, a scheme we call stochastically perturbed weights (SPW). We also ask whether it works, where and on which scales to inject the noise, and where it fails. A three-phase ablation across four deterministic backbones, Aurora, GraphCast, SFNO, and AIFS, selects one production baseline per model, benchmarked against the trained-probabilistic AIFS-ENS, FourCastNet 3 and Atlas as well as the operational ECMWF ensemble (IFS-ENS) over 112 initialisation times. At a 240 h (10-day) lead time the SPW ensembles reach continuous ranked probability skill scores (CRPSS) between 0.04 and 0.13 below the best trained-probabilistic baseline, at zero marginal training cost. No injection site works across models: the productive tensor group is architecture-specific, so SPW is at present a tuning procedure rather than a plug-and-play recipe. Its main failure mode is a coherent whole-field offset that overdisperses the domain mean, and restricting the noise to coarse scales or perturbing the initial conditions each repair part of it.

天气预测不确定性集合预报神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。