arXiv:2503.10907cs.MAcs.AI2025-03被引 6

H2-MARL用多智能体强化学习平衡疫情中医院压力与出行限制。

H2-MARL: Multi-Agent Reinforcement Learning for Pareto Optimality in Hospital Capacity Strain and Human Mobility during Epidemic

  • 将城市分区设为智能体,用双目标奖励函数协调防疫与出行。
  • 在四城超十亿记录数据上验证,可同时降低医院负荷和出行损失。
  • 适用于不同规模城市,适合智慧城市与公共政策研究者参考。

COVID-19后,如何在减少出行限制损失与保障医院容量之间取得有效平衡成为关注焦点。基于强化学习的人流管理策略虽能应对城市与疫情的动态演变,但在乡镇层级协同控制及跨规模适应性方面仍存挑战。为此,本文提出一种实现帕累托最优的多智能体强化学习方法(H2-MARL),适用于不同规模城市。首先构建具有在线更新参数的乡镇级感染模型,并建立城市级时空动态疫情模拟器;在此基础上,将每个行政区划视为智能体,设计兼顾医院容量与人流限制的双目标奖励函数,并引入专家知识增强经验回放缓冲区。为评估模型效果,构建了涵盖四个不同规模代表性城市的乡镇级人流数据集,包含超十亿条记录。大量实验表明,H2-MARL具备最优双目标权衡能力,在最小化医院容量压力的同时,最大限度减少出行限制损失;且其在不同规模城市中的适用性得到验证,展现出良好的实用性和通用性。

原文摘要 · Abstract (English)

The necessity of achieving an effective balance between minimizing the losses associated with restricting human mobility and ensuring hospital capacity has gained significant attention in the aftermath of COVID-19. Reinforcement learning (RL)-based strategies for human mobility management have recently advanced in addressing the dynamic evolution of cities and epidemics; however, they still face challenges in achieving coordinated control at the township level and adapting to cities of varying scales. To address the above issues, we propose a multi-agent RL approach that achieves Pareto optimality in managing hospital capacity and human mobility (H2-MARL), applicable across cities of different scales. We first develop a township-level infection model with online-updatable parameters to simulate disease transmission and construct a city-wide dynamic spatiotemporal epidemic simulator. On this basis, H2-MARL is designed to treat each division as an agent, with a trade-off dual-objective reward function formulated and an experience replay buffer enriched with expert knowledge built. To evaluate the effectiveness of the model, we construct a township-level human mobility dataset containing over one billion records from four representative cities of varying scales. Extensive experiments demonstrate that H2-MARL has the optimal dual-objective trade-off capability, which can minimize hospital capacity strain while minimizing human mobility restriction loss. Meanwhile, the applicability of the proposed model to epidemic control in cities of varying scales is verified, which showcases its feasibility and versatility in practical applications.

多智能体疫情管理强化学习城市规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。