用约束型强化学习优化城市检查调度,提升食品安全监管效率。
Optimizing Urban Service Allocation with Time-Constrained Restless Bandits
- 基于马尔可夫决策过程重构与整数规划前瞻,解决带时间窗的多臂老虎机难题
- 模拟与真实数据均显示检查效果提升24%~33%,且对突发检查有强鲁棒性
- 适用于公共健康、城市管理等需平衡合规与成本的智能调度场景
市政检查是保障商品和服务质量的重要手段。本文以芝加哥食品场所检查为例,研究如何智能调度检查任务以最大化其影响。芝加哥公共卫生局(CDPH)每年检查数千家场所,2023年有超过3,000份检查报告不合格。为平衡遵守规范、减少对商户干扰和控制检查成本,CDPH每年为每家场所设定一个检查窗口,并保证在该窗口内仅检查一次;同时承诺对突发食品安全事件或投诉进行突击检查。这些约束使得传统的无休止多臂老虎机(RMAB)方法难以适用。我们提出一种基于威特指数的扩展方法,可保证动作窗口约束与频率,并能优化窗口分配本身。核心思想是结合马尔可夫决策过程重构与基于整数规划的前瞻策略,以在约束下最大化检查成效。我们还构建了一个基于神经网络的监督学习模型,利用公开的CDPH检查记录建模真实芝加哥场所的状态转移,相比直接预测失败率提升了10% AUC。实验表明,本方法在模拟中实现最高24%的性能提升,在真实数据上达33%,且对突击检查具有强鲁棒性,同时揭示了调度约束的实际影响。
原文摘要 · Abstract (English)
Municipal inspections are an important part of maintaining the quality of goods and services. In this paper, we approach the problem of intelligently scheduling service inspections to maximize their impact, using the case of food establishment inspections in Chicago as a case study. The Chicago Department of Public Health (CDPH) inspects thousands of establishments each year, with a substantial fail rate (over 3,000 failed inspection reports in 2023). To balance the objectives of ensuring adherence to guidelines, minimizing disruption to establishments, and minimizing inspection costs, CDPH assigns each establishment an inspection window every year and guarantees that they will be inspected exactly once during that window. Meanwhile, CDPH also promises surprise public health inspections for unexpected food safety emergencies or complaints. These constraints create a challenge for a restless multi-armed bandit (RMAB) approach, for which there are no existing methods. We develop an extension to Whittle index-based systems for RMABs that can guarantee action window constraints and frequencies, and furthermore can be leveraged to optimize action window assignments themselves. Briefly, we combine MDP reformulation and integer programming-based lookahead to maximize the impact of inspections subject to constraints. A neural network-based supervised learning model is developed to model state transitions of real Chicago establishments using public CDPH inspection records, which demonstrates 10% AUC improvements compared with directly predicting establishments' failures. Our experiments not only show up to 24% (in simulation) or 33% (on real data) objective improvements resulting from our approach and robustness to surprise inspections, but also give insight into the impact of scheduling constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。