用分层强化学习优化疫情多集群防控资源分配。
Optimizing Resource-Constrained Non-Pharmaceutical Interventions for Multi-Cluster Outbreak Control Using Hierarchical Reinforcement Learning
- 分层架构:全局控制器调资源总量,局部策略评估个体投入价值。
- 实测提升20%-30%控制效果,40个集群并发时仍高效决策。
- 适合资源有限、多点爆发的公共卫生应急场景。
非药物干预(如检测、隔离)对控制传染病暴发至关重要,但常受资源限制,尤其在疫情初期。现实中,需在多个异步出现、规模与风险各异的疫情集群间分配共享资源。每个集群由单一感染者关联的密接者构成,决策需在不确定性与差异化需求下进行,并满足操作约束。本文将问题建模为受限的随机多臂老虎机(RMAB),提出分层强化学习框架:全局控制器学习连续动作成本倍增因子以调节总资源需求,局部策略估算各集群内个体资源投入的边际价值。在基于SARS-CoV-2的代理模拟器中评估,涵盖多种系统规模与检测预算,该方法显著优于RMAB启发式及启发式基线,控制效果提升20%-30%。实验验证了最多40个并发活跃集群下的可扩展性,决策速度超越传统RMAB方法。
原文摘要 · Abstract (English)
Non-pharmaceutical interventions (NPIs), such as diagnostic testing and quarantine, are crucial for controlling infectious disease outbreaks but are often constrained by limited resources, particularly in early outbreak stages. In real-world public health settings, resources must be allocated across multiple outbreak clusters that emerge asynchronously, vary in size and risk, and compete for a shared resource budget. Here, a cluster corresponds to a group of close contacts generated by a single infected index case. Thus, decisions must be made under uncertainty and heterogeneous demands, while respecting operational constraints. We formulate this problem as a constrained restless multi-armed bandit and propose a hierarchical reinforcement learning framework. A global controller learns a continuous action cost multiplier that adjusts global resource demand, while a generalized local policy estimates the marginal value of allocating resources to individuals within each cluster. We evaluate the proposed framework in a realistic agent-based simulator of SARS-CoV-2 with dynamically arriving clusters. Across a wide range of system scales and testing budgets, our method consistently outperforms RMAB-inspired and heuristic baselines, improving outbreak control effectiveness by 20%-30%. Experiments on up to 40 concurrently active clusters further demonstrate that the hierarchical framework is highly scalable and enables faster decision-making than the RMAB-inspired method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。