arXiv:2409.02246cs.LGmath.OC2024-09被引 6

用多智能体强化学习联合优化警力巡逻与出警,提升响应速度

Multi-Agent Reinforcement Learning for Joint Police Patrol and Dispatch

  • 将每个警员视为独立智能体,共享深度Q网络学习联合策略
  • 联合优化后响应时间显著优于单独优化巡逻或出警的策略
  • 适合关注警务效率与公平性的城市管理者和政策制定者

警察巡逻单位需在预防性巡逻和应急事件处置之间分配时间。现有研究通常将巡逻与出警决策分开处理。本文提出一种联合优化方法,以提升警务运作效率并缩短应急呼叫响应时间。我们设计了一种异构多智能体强化学习框架,将每位巡警视为独立的Q-learner,共享一个深度Q网络来表示状态-动作价值。出警决策通过混合整数规划与组合动作空间的价值函数近似进行选择。实验表明,该方法能够学习到优于仅针对巡逻或出警单独优化的联合策略。管理启示:联合优化的策略可在保障效率的同时灵活实现公平性等目标。

原文摘要 · Abstract (English)

Police patrol units need to split their time between performing preventive patrol and being dispatched to serve emergency incidents. In the existing literature, patrol and dispatch decisions are often studied separately. We consider joint optimization of these two decisions to improve police operations efficiency and reduce response time to emergency calls. Methodology/results: We propose a novel method for jointly optimizing multi-agent patrol and dispatch to learn policies yielding rapid response times. Our method treats each patroller as an independent Q-learner (agent) with a shared deep Q-network that represents the state-action values. The dispatching decisions are chosen using mixed-integer programming and value function approximation from combinatorial action spaces. We demonstrate that this heterogeneous multi-agent reinforcement learning approach is capable of learning joint policies that outperform those optimized for patrol or dispatch alone. Managerial Implications: Policies jointly optimized for patrol and dispatch can lead to more effective service while targeting demonstrably flexible objectives, such as those encouraging efficiency and equity in response.

多智能体强化学习警务优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。