arXiv:2501.08020cs.AI2025-01被引 4

用多智能体强化学习优化城市巡警路线,提升犯罪高发区覆盖率。

Cooperative Patrol Routing: Optimizing Urban Crime Surveillance through Multi-Agent Reinforcement Learning

  • 基于去中心化马尔可夫决策过程设计协同巡逻策略。
  • 在马拉加三区测试中,对3%高犯罪节点覆盖超90%,20%节点覆盖达65%。
  • 提出新评估指标覆盖指数,适合警务规划与智能调度研究者。

有效设计巡逻策略是一项复杂挑战,尤其在中大型区域。本文提出一种基于去中心化部分可观测马尔可夫决策过程的多智能体强化学习(MARL)模型,用于在无向图表示的城市环境中规划难以预测的巡逻路线。目标是最大化特定时间窗内环境的覆盖程度。模型在西班牙马拉加市三个中等规模行政区进行测试,依据真实犯罪数据优化警力巡逻路径,以最大化高犯罪风险区域的监控覆盖。比较多种MARL算法后,价值分解近端策略优化(VDPPO)表现最佳。我们引入新型评价指标——覆盖指数,受犯罪学中预测准确指数(PAI)启发,用于评估不同场景下(智能体数量、起始位置、观测范围变化)的覆盖性能。结果表明,协调生成的路线对占比3%的高犯罪节点覆盖超过90%,对20%节点覆盖达65%,符合警方资源分配标准。

原文摘要 · Abstract (English)

The effective design of patrol strategies is a difficult and complex problem, especially in medium and large areas. The objective is to plan, in a coordinated manner, the optimal routes for a set of patrols in a given area, in order to achieve maximum coverage of the area, while also trying to minimize the number of patrols. In this paper, we propose a multi-agent reinforcement learning (MARL) model, based on a decentralized partially observable Markov decision process, to plan unpredictable patrol routes within an urban environment represented as an undirected graph. The model attempts to maximize a target function that characterizes the environment within a given time frame. Our model has been tested to optimize police patrol routes in three medium-sized districts of the city of Malaga. The aim was to maximize surveillance coverage of the most crime-prone areas, based on actual crime data in the city. To address this problem, several MARL algorithms have been studied, and among these the Value Decomposition Proximal Policy Optimization (VDPPO) algorithm exhibited the best performance. We also introduce a novel metric, the coverage index, for the evaluation of the coverage performance of the routes generated by our model. This metric is inspired by the predictive accuracy index (PAI), which is commonly used in criminology to detect hotspots. Using this metric, we have evaluated the model under various scenarios in which the number of agents (or patrols), their starting positions, and the level of information they can observe in the environment have been modified. Results show that the coordinated routes generated by our model achieve a coverage of more than $90\%$ of the $3\%$ of graph nodes with the highest crime incidence, and $65\%$ for $20\%$ of these nodes; $3\%$ and $20\%$ represent the coverage standards for police resource allocation.

巡逻优化强化学习城市安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。