arXiv:2606.04562cs.AIcs.LG2026-06

考虑行为与执行误差,用强化学习优化防疫政策。

Neetyabhas: A Framework for Uncertainty-Aware Public Policy Optimization in Rational Agent-Based Models

论文配图:Neetyabhas: A Framework for Uncertainty-Aware Public Policy Optimization in Rational Agent-Based Models
图 1 · 摘自论文原文
  • 用分层强化学习模拟1000人实时决策与政策执行。
  • 戴口罩和接种疫苗显著降低疫情峰值与持续时间。
  • 适合研究公共卫生政策设计与不确定性应对的人看。

世界卫生组织的新冠非药物干预措施(如封锁、疫苗接种)虽能有效控制传播,但带来沉重经济负担。现有研究常忽视个体行为,错误假设感染情况完全可追踪且政策执行无误,未能反映真实世界的不确定性和误差。我们提出一种整合疫情测量(感染/住院)和政策执行不确定性的综合方法。构建了包含1000名个体的仿真模型,他们实时决定佩戴口罩、接种疫苗和外出购物;同时,政策制定者基于健康与经济观测部署封锁、强制令等干预措施。该框架采用分层强化学习代理,结合深度Q网络及不确定性感知的策略梯度变体(DDPG和TD3)。模拟结果表明,疫情进程得到有效控制。口罩与疫苗被证明极为有效,显著降低了疫情峰值高度与持续时间。通过融合个体行为、政策不确定性与多维度干预,本动态调控方法成功减轻了疫情冲击。结论指出,将不确定性与人类行为纳入公共卫生政策框架,克服了以往研究局限。仿真显示,在复杂大流行中,考虑个体选择与不完美数据对设计有效干预至关重要,口罩与疫苗是关键工具。

原文摘要 · Abstract (English)

Purpose The WHO's COVID-19 non-pharmaceutical interventions (e.g., lockdowns, vaccinations) effectively curb transmission but impose heavy economic strains. Existing research often neglects individual behaviors and falsely assumes perfect infection tracking and flawless policy execution, failing to account for real-world uncertainties and errors. Methods We propose an integrative approach incorporating uncertainties in both epidemic measurement (infections/hospitalizations) and policy implementation. We built a simulation model of 1,000 individuals making real-time choices regarding mask-wearing, vaccination, and shopping. Concurrently, policymakers deploy interventions (lockdowns, mandates) based on health and economic observations. This framework is driven by hierarchical reinforcement learning agents, utilizing deep Q-networks alongside uncertainty-aware policy gradient variants (DDPG and TD3). Results The simulations effectively managed the epidemic's progression. Masking and vaccinations proved highly effective, significantly reducing both the outbreak's peak height and duration. By integrating individual behaviors, policy uncertainties, and multifaceted interventions, our dynamic control approach successfully mitigated the epidemic's impact. Conclusions Our model overcomes previous research limitations by embedding uncertainty and human behavior into public health policy frameworks. The simulation demonstrates that accounting for individual choices and imperfect data is crucial for designing effective interventions during complex pandemics, with masks and vaccines serving as pivotal tools.

公共政策强化学习疫情模拟不确定性建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。