用多模态大模型描述物体互动,让视频异常检测既准确又可解释。
Leveraging Multimodal LLM Descriptions of Activity for Explainable Semi-Supervised Video Anomaly Detection
- 通过查询多模态大模型生成物体互动的文本描述作为高层表征
- 在无交互异常数据集上达到当前最佳性能,复杂交互异常检测显著提升
- 天然具备可解释性,适用于需要透明决策的场景
现有半监督视频异常检测方法在涉及物体交互的复杂异常检测上表现不佳,且缺乏可解释性。为此,我们提出一种基于多模态大语言模型(MLLM)的新框架。不同于以往直接在帧级别做出异常判断的方法,本方法聚焦于提取和解释物体随时间的活动与交互。通过向MLLM输入不同时间点的物体对视觉信息,生成来自正常视频的活动与交互文本描述。这些文本描述作为视频中物体行为的高层表征,在测试时通过与训练阶段正常视频的描述对比来检测异常。该方法天然具备可解释性,并可与多种传统VAD方法结合以增强其可解释性。在多个基准数据集上的大量实验表明,该方法不仅能有效检测基于交互的复杂异常,还在无交互异常的数据集上达到最先进水平。
原文摘要 · Abstract (English)
Existing semi-supervised video anomaly detection (VAD) methods often struggle with detecting complex anomalies involving object interactions and generally lack explainability. To overcome these limitations, we propose a novel VAD framework leveraging Multimodal Large Language Models (MLLMs). Unlike previous MLLM-based approaches that make direct anomaly judgments at the frame level, our method focuses on extracting and interpreting object activity and interactions over time. By querying an MLLM with visual inputs of object pairs at different moments, we generate textual descriptions of the activity and interactions from nominal videos. These textual descriptions serve as a high-level representation of the activity and interactions of objects in a video. They are used to detect anomalies during test time by comparing them to textual descriptions found in nominal training videos. Our approach inherently provides explainability and can be combined with many traditional VAD methods to further enhance their interpretability. Extensive experiments on benchmark datasets demonstrate that our method not only detects complex interaction-based anomalies effectively but also achieves state-of-the-art performance on datasets without interaction anomalies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。