arXiv:2609.02566cs.LG2026-09

用在线强化学习修正天气模型,提升预报精度且保持稳定。

Online Reinforcement Learning in the Met Office Unified Model through Distributed Model-Agent Coupling

论文配图:Online Reinforcement Learning in the Met Office Unified Model through Distributed Model-Agent Coupling
图 1 · 摘自论文原文
  • 通过分布式强化学习代理与气象模型耦合,实时学习温度修正策略。
  • 在六个纬度带中,500百帕高度误差降低最多达45.8%,海平面气压误差降27.3%。
  • 适合关注气象模型智能修正、可解释性强化学习的科研与业务人员。

机器学习修正可补充数值天气预报,前提是能适应模型状态演变并保持动力一致性与数值稳定性。为在全局预报模型中验证这一设想,我们通过秩局部张量将英国气象局(UKMO)统一模式(UM)与分布式强化学习代理耦合。采用DDPG算法的演员网络在每个大气柱的70个垂直层间共享权重,并对模型倾向施加有界位温修正。在十次受控训练预报中,通过向英国气象局运行分析场逼近来提供即时反事实目标。随后冻结策略,在非受控预报中进行推理评估。耦合工作流成功完成训练,且在测试案例中保持数值稳定。相对于匹配的原生UM预报在+6小时时,学习到的策略在六个纬度带中的四个降低了Z$_{500}$ MAE,其中南北热带分别减少45.8%和40.8%;海平面气压误差在三个纬度带下降,最大降幅达27.3%(0–30°N)。此单例实验展示了分布式在线学习结合非受控推理的显著前景与可行性,为操作型系统中基于强化学习的偏差修正与参数化奠定了基础。

原文摘要 · Abstract (English)

Machine-learnt corrections can complement numerical weather prediction only if they adapt to the evolving model state while preserving dynamical consistency and numerical stability. To test this within a global forecasting model, we couple the Met Office (UKMO) Unified Model (UM) with distributed RL agents through rank-local tensors. A DDPG actor shares weights across the 70 vertical model levels of each atmospheric column and applies bounded potential-temperature corrections to the model tendencies. Across ten nudged training forecasts, nudging calculations towards the UKMO operational analysis provides an immediate counterfactual target. The frozen policy is then evaluated in a non-nudged forecast for inference. The coupled workflow successfully completes training and remains numerically stable in the evaluated case. Relative to a matched native UM forecast at +6 h, the learnt policy reduces Z$_{500}$ MAE in four of six latitude bands, including reductions of 45.8% and 40.8% in the northern and southern tropics. MSLP error too decreases in three bands, with a maximum reduction of 27.3% at 0-30{\deg}N. This single-case experiment demonstrates significant promise and feasibility of distributed online learning followed by non-nudged inference, laying the groundwork for RL-based bias correction and parametrisations within operational systems.

天气预报强化学习模型修正数值模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。