为强化学习在疫情模拟中的应用设计领域驱动的评估指标。
Domain-driven Metrics for Reinforcement Learning: A Case Study on Epidemic Control using Agent-based Simulation

- 基于疫情传播机制设计专属奖励函数,提升算法合理性
- 在不同口罩供应情景下验证策略优化效果,提升决策有效性
- 适合从事社会仿真与公共政策建模的研究者参考
在代理模型(ABMs)和理性代理模型(RABMs)的开发与优化中,强化学习被广泛应用。然而,由于系统复杂性和随机性,以及缺乏标准化的评估指标,评估基于强化学习的模型性能仍具挑战。本研究提出一种领域驱动的强化学习评估指标体系,在理性代理模型疾病传播案例中,用于模拟戴口罩、接种疫苗和封锁等行为。通过在不同口罩可得性情景下进行策略优化,验证了领域驱动奖励与传统及先进指标结合的有效性,显著提升了模型对现实政策干预的响应能力。
原文摘要 · Abstract (English)
For the development and optimization of agent-based models (ABMs) and rational agent-based models (RABMs), optimization algorithms such as reinforcement learning are extensively used. However, assessing the performance of RL-based ABMs and RABMS models is challenging due to the complexity and stochasticity of the modeled systems, and the lack of well-standardized metrics for comparing RL algorithms. In this study, we are developing domain-driven metrics for RL, while building on state-of-the-art metrics. We demonstrate our ``Domain-driven-RL-metrics'' using policy optimization on a rational ABM disease modeling case study to model masking behavior, vaccination, and lockdown in a pandemic. Our results show the use of domain-driven rewards in conjunction with traditional and state-of-the-art metrics for a few different simulation scenarios such as the differential availability of masks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。