arXiv:2511.18035stat.MEcs.LG2025-11被引 1

用强化学习动态调整防疫措施,降低重症负担同时减少社会成本

On a Reinforcement Learning Methodology for Epidemic Control, with application to COVID-19

  • 结合流行病模型与贝叶斯推理,用强化学习自适应选择干预强度
  • 相比历史政策,在300天内显著降低重症床位压力
  • 适合政策制定者和公共卫生研究者参考决策支持系统设计

本文提出一种实时、数据驱动的传染病控制决策支持框架。将分室流行病模型与序贯贝叶斯推断及强化学习(RL)控制器相结合,自适应地选择干预水平,以平衡疾病负担(如重症监护病房负荷)与社会经济成本。通过实证实验与专家反馈构建特定情境的成本函数。研究两种RL策略:基于蒙特卡洛网格搜索计算的重症阈值规则,以及基于后验平均Q学习代理的策略。通过拟合英国新冠疫情期间的公共重症监护病房占用数据,对每种RL控制器生成反事实推演情景,从而对比其与历史政府策略的效果。在300天周期内,针对多种成本参数,两种控制器均显著降低重症负担,证明了贝叶斯序贯学习与强化学习结合在辅助疫情管控政策设计中的有效性。

原文摘要 · Abstract (English)

This paper presents a real time, data driven decision support framework for epidemic control. We combine a compartmental epidemic model with sequential Bayesian inference and reinforcement learning (RL) controllers that adaptively choose intervention levels to balance disease burden, such as intensive care unit (ICU) load, against socio economic costs. We construct a context specific cost function using empirical experiments and expert feedback. We study two RL policies: an ICU threshold rule computed via Monte Carlo grid search, and a policy based on a posterior averaged Q learning agent. We validate the framework by fitting the epidemic model to publicly available ICU occupancy data from the COVID 19 pandemic in England and then generating counterfactual roll out scenarios under each RL controller, which allows us to compare the RL policies to the historical government strategy. Over a 300 day period and for a range of cost parameters, both controllers substantially reduce ICU burden relative to the observed interventions, illustrating how Bayesian sequential learning combined with RL can support the design of epidemic control policies.

强化学习疫情控制决策支持

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。