arXiv:2508.06189cs.CV2025-08

通过多智能体异步协作,实时预测公共场景犯罪行为。

MA-CBP: A Criminal Behavior Prediction Framework Based on Multi-Agent Asynchronous Collaboration

  • 构建多智能体系统,异步处理视频帧与历史语义,实现跨时序推理。
  • 在多个数据集上超越现有方法,实现早期犯罪风险预警。
  • 适用于城市公共安全监控,适合需要实时行为分析的场景。

随着城市化进程加速,公共场景中的犯罪行为对社会安全构成日益严重的威胁。基于特征识别的传统异常检测方法难以从历史信息中捕捉高层行为语义,而基于大语言模型的生成方法通常无法满足实时性要求。为此,我们提出MA-CBP框架,一种基于多智能体异步协作的犯罪行为预测方法。该框架将实时视频流转化为帧级语义描述,构建因果一致的历史摘要,并融合相邻图像帧,实现对长短期上下文的联合推理。生成的行为决策包含事件主体、地点和原因等关键要素,支持潜在犯罪活动的早期预警。此外,我们构建了一个高质量的犯罪行为数据集,提供多尺度语言监督,涵盖帧级、摘要级和事件级语义标注。实验结果表明,该方法在多个数据集上均取得优异性能,为城市公共安全风险预警提供了有前景的解决方案。

原文摘要 · Abstract (English)

With the acceleration of urbanization, criminal behavior in public scenes poses an increasingly serious threat to social security. Traditional anomaly detection methods based on feature recognition struggle to capture high-level behavioral semantics from historical information, while generative approaches based on Large Language Models (LLMs) often fail to meet real-time requirements. To address these challenges, we propose MA-CBP, a criminal behavior prediction framework based on multi-agent asynchronous collaboration. This framework transforms real-time video streams into frame-level semantic descriptions, constructs causally consistent historical summaries, and fuses adjacent image frames to perform joint reasoning over long- and short-term contexts. The resulting behavioral decisions include key elements such as event subjects, locations, and causes, enabling early warning of potential criminal activity. In addition, we construct a high-quality criminal behavior dataset that provides multi-scale language supervision, including frame-level, summary-level, and event-level semantic annotations. Experimental results demonstrate that our method achieves superior performance on multiple datasets and offers a promising solution for risk warning in urban public safety scenarios.

犯罪预测多智能体视频理解实时推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。