arXiv:2510.07971cs.LGcs.MA2025-10被引 1

用高效代理模型让多智能体强化学习更快更准地模拟气候政策。

Climate Surrogates for Scalable Multi-Agent Reinforcement Learning: A Case Study with CICERO-SCM

  • 用预训练的循环神经网络替代复杂气候模型,嵌入强化学习环境。
  • 温度预测误差仅0.0004K,推理速度提升1000倍,训练加速超100倍。
  • 适合研究气候政策、大规模多智能体模拟的科研人员使用。

气候政策研究需要能反映多种温室气体对全球气温综合影响的模型,但这类模型计算成本高,难以嵌入强化学习框架。本文提出一种多智能体强化学习(MARL)框架,将高保真、高效率的气候代理模型直接集成到环境循环中,使区域智能体能在多气体动态下学习气候政策。作为概念验证,我们引入一个循环神经网络架构,在20,000条多气体排放路径上预训练,以替代气候模型CICERO-SCM。该代理模型达到接近模拟器的精度,全球平均温度均方根误差约0.0004 K,单步推理速度提升约1000倍。在气候政策MARL设置中替换原模拟器后,端到端训练速度提升超过100倍。我们证明代理模型与模拟器收敛至相同最优策略,并提出一种在模拟器不可行时评估该性质的方法。本工作绕过核心计算瓶颈,同时保持政策保真度,使大规模多智能体实验在多气体动态和高保真气候响应下成为可能。

原文摘要 · Abstract (English)

Climate policy studies require models that capture the combined effects of multiple greenhouse gases on global temperature, but these models are computationally expensive and difficult to embed in reinforcement learning. We present a multi-agent reinforcement learning (MARL) framework that integrates a high-fidelity, highly efficient climate surrogate directly in the environment loop, enabling regional agents to learn climate policies under multi-gas dynamics. As a proof of concept, we introduce a recurrent neural network architecture pretrained on ($20{,}000$) multi-gas emission pathways to surrogate the climate model CICERO-SCM. The surrogate model attains near-simulator accuracy with global-mean temperature RMSE $\approx 0.0004 \mathrm{K}$ and approximately $1000\times$ faster one-step inference. When substituted for the original simulator in a climate-policy MARL setting, it accelerates end-to-end training by $>\!100\times$. We show that the surrogate and simulator converge to the same optimal policies and propose a methodology to assess this property in cases where using the simulator is intractable. Our work allows to bypass the core computational bottleneck without sacrificing policy fidelity, enabling large-scale multi-agent experiments across alternative climate-policy regimes with multi-gas dynamics and high-fidelity climate response.

多智能体气候模拟强化学习代理模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。