arXiv:2510.03340cs.LGcs.AI2025-10

用多目标强化学习优化防疫政策,平衡防控与经济影响。

Learning Pareto-Optimal Pandemic Intervention Policies with MORL

  • 基于多目标强化学习与高精度疫情模拟器设计干预策略
  • 揭示疫情控制与经济稳定间的直接权衡关系
  • 可推广至多种传染病,支持透明化政策决策

COVID-19大流行凸显了在疾病控制与经济社会稳定之间寻求平衡的迫切需求。本文提出一种基于多目标强化学习(MORL)的框架,结合全新校准并验证于全球新冠数据的随机微分方程(SDE)疫情模拟器,以建模和评估防疫策略。该模拟器在国家尺度上的疫情动态重现精度较现有强化学习常用模型高出数个数量级。在该模拟器上训练的帕累托条件网络(PCN)代理,直观展示了新冠防控与经济稳定之间的直接权衡。此外,我们证明该框架具有普适性,可扩展至脊髓灰质炎、流感等不同传播特征的病原体,并发现其诱导出根本不同的干预策略。为贴近现实政策挑战,我们将模型应用于麻疹暴发,量化显示仅5%的疫苗覆盖率下降即需更严格、更昂贵的干预措施才能遏制传播。本工作提供了一个稳健且可适应的框架,支持公共健康危机中的透明、证据驱动型决策。

原文摘要 · Abstract (English)

The COVID-19 pandemic underscored a critical need for intervention strategies that balance disease containment with socioeconomic stability. We approach this challenge by designing a framework for modeling and evaluating disease-spread prevention strategies. Our framework leverages multi-objective reinforcement learning (MORL) - a formulation necessitated by competing objectives - combined with a new stochastic differential equation (SDE) pandemic simulator, calibrated and validated against global COVID-19 data. Our simulator reproduces national-scale pandemic dynamics with orders of magnitude higher fidelity than other models commonly used in reinforcement learning (RL) approaches to pandemic intervention. Training a Pareto-Conditioned Network (PCN) agent on this simulator, we illustrate the direct policy trade-offs between epidemiological control and economic stability for COVID-19. Furthermore, we demonstrate the framework's generality by extending it to pathogens with different epidemiological profiles, such as polio and influenza, and show how these profiles lead the agent to discover fundamentally different intervention policies. To ground our work in contemporary policymaking challenges, we apply the model to measles outbreaks, quantifying how a modest 5% drop in vaccination coverage necessitates significantly more stringent and costly interventions to curb disease spread. This work provides a robust and adaptable framework to support transparent, evidence-based policymaking for mitigating public health crises.

多目标学习疫情模拟政策优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。