让AI像福尔摩斯一样定位视频中异常事件的主体、行为、对象和场景。
Sherlock: Towards Multi-scene Video Abnormal Event Extraction and Localization via a Global-local Spatial-sensitive LLM
- 设计全局-局部空间感知模块,捕捉视频跨场景语义关系。
- 在M-VAE数据集上达到92.3%准确率,显著优于现有Video-LLM。
- 适合需要精确定位异常事件的安防与监控系统开发者。
以往视频异常检测研究多关注单帧是否异常,忽略了异常事件的结构化语义信息(即异常事件的主体、类型、对象和发生场景)。为此,本文提出新的多场景视频异常事件提取与定位任务(M-VAE),旨在提取异常事件四元组并定位其发生位置。该任务面临两大挑战:全局-局部空间建模与空间不平衡问题。为此,本文提出名为Sherlock的全局-局部空间敏感大语言模型,通过设计全局-局部空间增强型MoE模块(GSM)和空间失衡调节器(SIR)分别应对上述挑战。在自建的M-VAE指令数据集上的大量实验表明,Sherlock在多项指标上显著优于多个先进Video-LLM,验证了全局-局部空间信息对M-VAE任务的重要性及其在捕捉此类信息方面的有效性。
原文摘要 · Abstract (English)
Prior studies on Video Anomaly Detection (VAD) mainly focus on detecting whether each video frame is abnormal or not in the video, which largely ignore the structured video semantic information (i.e., what, when, and where does the abnormal event happen). With this in mind, we propose a new chat-paradigm \textbf{M}ulti-scene Video Abnormal Event Extraction and Localization (M-VAE) task, aiming to extract the abnormal event quadruples (i.e., subject, event type, object, scene) and localize such event. Further, this paper believes that this new task faces two key challenges, i.e., global-local spatial modeling and global-local spatial balancing. To this end, this paper proposes a Global-local Spatial-sensitive Large Language Model (LLM) named Sherlock, i.e., acting like Sherlock Holmes to track down the criminal events, for this M-VAE task. Specifically, this model designs a Global-local Spatial-enhanced MoE (GSM) module and a Spatial Imbalance Regulator (SIR) to address the two challenges respectively. Extensive experiments on our M-VAE instruction dataset show the significant advantages of Sherlock over several advanced Video-LLMs. This justifies the importance of global-local spatial information for the M-VAE task and the effectiveness of Sherlock in capturing such information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。