arXiv:2601.04268cs.LGphysics.ao-ph2026-01被引 1

用强化学习让气候模型自动调节参数,提升预测准确性。

Replacing Tunable Parameters in Weather and Climate Models with State-Dependent Functions using Reinforcement Learning

  • 用强化学习在线学习状态依赖的参数函数,替代固定系数。
  • 在理想化模型中,平均误差降低,热带和中纬度表现最优。
  • 方法物理可解释,适合想改进气候模型的科研人员。

天气和气候模型依赖参数化方案来表示未解析的亚网格过程。传统方法使用固定系数,约束弱且离线调优,导致持续偏差,限制模型对底层物理的适应能力。本研究提出一种框架,利用强化学习(RL)在线学习参数化方案的组件,使其随模型状态动态变化,并在三个理想化测试平台(简单气候偏差修正SCBC、辐射-对流平衡RCE、单层能量平衡模型EBM)上评估策略驱动的参数更新,涵盖单智能体与联邦多智能体设置。九种RL算法中,截断分位数评论者(TQC)、深度确定性策略梯度(DDPG)和双延迟深度确定性策略梯度(TD3)表现最佳,收敛稳定,性能通过加权均方根误差(area-weighted RMSE)及温度/气压层诊断对比静态基线。在EBM中,单智能体RL优于静态调参,热带与中纬度改善显著;联邦多智能体设置实现专业化控制与更快收敛,六智能体DDPG频繁聚合配置在热带与中纬度取得最低加权误差。学习到的修正具有物理意义:智能体调节EBM辐射参数以减少经向偏差,调整RCE递减率匹配垂直温度误差,稳定加热增量以抑制漂移。结果表明,强化学习可在理想化场景中学习有效状态依赖的参数化组件,为数值模型中的在线学习提供可扩展路径,并为天气与气候模型评估奠定基础。

原文摘要 · Abstract (English)

Weather and climate models rely on parametrisations to represent unresolved sub-grid processes. Traditional schemes rely on fixed coefficients that are weakly constrained and tuned offline, contributing to persistent biases that limit their ability to adapt to underlying physics. This study presents a framework that learns components of parametrisation schemes online as a function of the evolving model state using reinforcement learning (RL) and evaluates policy-driven parameter updates across idealised testbeds spanning a simple climate bias correction (SCBC), a radiative-convective equilibrium (RCE), and a zonal mean energy balance model (EBM) with single-agent and federated multi-agent settings. Across nine RL algorithms, Truncated Quantile Critics (TQC), Deep Deterministic Policy Gradient (DDPG), and Twin Delayed DDPG (TD3) achieved the highest skill and stable convergence, with performance assessed against a static baseline using area-weighted RMSE, temperature and pressure-level diagnostics. For the EBM, single-agent RL outperformed static parameter tuning with the strongest gains in tropical and mid-latitude bands, while federated RL on multi-agent setups enabled specialised control and faster convergence, with a six-agent DDPG configuration using frequent aggregation yielding the lowest area-weighted RMSE across the tropics and mid-latitudes. The learnt corrections were also physically meaningful as agents modulated EBM radiative parameters to reduce meridional biases, adjusted RCE lapse rates to match vertical temperature errors, and stabilised heating increments to limit drift. Overall, results show that RL can learn skilful state-dependent parametrisation components in idealised settings, offering a scalable pathway for online learning within numerical models and a starting point for evaluation in weather and climate models.

气候模型强化学习参数化在线学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。