arXiv:2410.07287physics.soc-phcs.AI2024-10被引 7

用强化学习模拟各国合作与竞争,探索更优气候路径。

Crafting Desirable Climate Trajectories with RL Explored Socio-Environmental Simulations

  • 多智能体强化学习模拟不同国家/利益方的互动决策
  • 合作时能稳定达成低碳且经济改善的未来路径
  • 竞争环境下多数情况难以实现理想气候目标

气候变化构成生存威胁,亟需有效政策推动变革。决策复杂,涉及多方冲突与证据不确定性。近年来,政策制定者越来越多地借助仿真与计算方法,其中综合评估模型(IAMs)融合社会、经济与环境模拟以预测政策效果,联合国在最新政府间气候变化专门委员会(IPCC)报告中即采用其输出。传统方法依赖递归方程求解器,但在不确定性下表现不佳。近期初步研究显示,用强化学习(RL)替代传统求解器,在不确定与噪声场景中表现更优。本文进一步引入多个交互式RL智能体,初步分析驱动当前气候危机的多方社会互动机制。结果表明:合作智能体可持续规划出碳排放更低、经济更优的未来路径;但引入竞争(如对立奖励函数)后,理想气候未来极少达成。为提升仿真现实性,我们通过可视化状态空间识别导致行为不稳定的区域,以理解算法失效原因。最后,指出现有局限及未来改进方向,确保该技术可用于政策推导。

原文摘要 · Abstract (English)

Climate change poses an existential threat, necessitating effective climate policies to enact impactful change. Decisions in this domain are incredibly complex, involving conflicting entities and evidence. In the last decades, policymakers increasingly use simulations and computational methods to guide some of their decisions. Integrated Assessment Models (IAMs) are one of such methods, which combine social, economic, and environmental simulations to forecast potential policy effects. For example, the UN uses outputs of IAMs for their recent Intergovernmental Panel on Climate Change (IPCC) reports. Traditionally these have been solved using recursive equation solvers, but have several shortcomings, e.g. struggling at decision making under uncertainty. Recent preliminary work using Reinforcement Learning (RL) to replace the traditional solvers shows promising results in decision making in uncertain and noisy scenarios. We extend on this work by introducing multiple interacting RL agents as a preliminary analysis on modelling the complex interplay of socio-interactions between various stakeholders or nations that drives much of the current climate crisis. Our findings show that cooperative agents in this framework can consistently chart pathways towards more desirable futures in terms of reduced carbon emissions and improved economy. However, upon introducing competition between agents, for instance by using opposing reward functions, desirable climate futures are rarely reached. Modelling competition is key to increased realism in these simulations, as such we employ policy interpretation by visualising what states lead to more uncertain behaviour, to understand algorithm failure. Finally, we highlight the current limitations and avenues for further work to ensure future technology uptake for policy derivation.

强化学习气候政策多智能体仿真建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。