首个支持实时视频异常预测、检测与分析的智能助手
AssistPDA: An Online Video Surveillance Assistant for Video Anomaly Prediction, Detection, and Analysis
- 构建统一框架实现流式视频的实时异常预测与检测
- 提出事件级异常预测,可提前发现潜在异常行为
- 适配实际安防场景,支持用户交互与长期时序建模
大语言模型(LLM)的发展推动了基于LLM的视频异常检测(VAD)研究,但现有方法多聚焦于视频级异常问答或离线检测,忽视了实际应用中对实时性的要求。为弥合这一差距并促进基于LLM的VAD落地,我们提出AssistPDA,首个统一视频异常预测、检测与分析(VAPDA)的在线视频监控助手。AssistPDA支持流式视频的实时推理,并具备用户交互能力。我们引入新的事件级异常预测任务,可在异常完全发生前进行前瞻性预警。为增强对复杂时空关系的建模能力,提出时空关系蒸馏(STRD)模块,将视觉-语言模型(VLM)在离线场景下的长时序建模能力迁移至实时环境,赋予系统强时空依赖理解与长序列记忆能力。此外,我们构建了首个面向基于VLM的在线VAPDA的大规模基准数据集VAPDA-127K。大量实验表明,AssistPDA超越现有离线VLM方法,在实时VAPDA上达到新基准。相关数据集与代码将开源,以推动社区研究。
原文摘要 · Abstract (English)
The rapid advancements in large language models (LLMs) have spurred growing interest in LLM-based video anomaly detection (VAD). However, existing approaches predominantly focus on video-level anomaly question answering or offline detection, ignoring the real-time nature essential for practical VAD applications. To bridge this gap and facilitate the practical deployment of LLM-based VAD, we introduce AssistPDA, the first online video anomaly surveillance assistant that unifies video anomaly prediction, detection, and analysis (VAPDA) within a single framework. AssistPDA enables real-time inference on streaming videos while supporting interactive user engagement. Notably, we introduce a novel event-level anomaly prediction task, enabling proactive anomaly forecasting before anomalies fully unfold. To enhance the ability to model intricate spatiotemporal relationships in anomaly events, we propose a Spatio-Temporal Relation Distillation (STRD) module. STRD transfers the long-term spatiotemporal modeling capabilities of vision-language models (VLMs) from offline settings to real-time scenarios. Thus it equips AssistPDA with a robust understanding of complex temporal dependencies and long-sequence memory. Additionally, we construct VAPDA-127K, the first large-scale benchmark designed for VLM-based online VAPDA. Extensive experiments demonstrate that AssistPDA outperforms existing offline VLM-based approaches, setting a new state-of-the-art for real-time VAPDA. Our dataset and code will be open-sourced to facilitate further research in the community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。