让AI像人一样一步步推理视频异常,提升理解深度。
Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought
- 构建感知到认知的思维链,引导模型分步推理异常
- 在新数据集上超越现有模型,在异常检测与推理任务均领先
- 适合关注AI可解释性与复杂视觉理解的研究者
多模态大语言模型(MLLM)在复杂视觉任务中展现出强大推理能力。然而,现有基于MLLM的视频异常检测(VAD)方法仍停留在浅层描述,缺乏深层推理。本文提出新任务视频异常推理(VAR),旨在通过要求MLLM在回答前显式思考,实现对视频异常的深度分析与理解。为此,我们设计了感知到认知的思维链(P2C-CoT),模拟人类识别异常的过程,引导模型逐步推理。基于此结构,我们构建了专用数据集Vad-Reasoning。此外,提出改进的强化学习算法AVA-GRPO,通过自验证机制,在有限标注下显式激励模型的推理能力。实验表明,Vad-R1在VAD和VAR任务上均优于开源与专有模型。代码与数据集将公开于https://github.com/wbfwonderful/Vad-R1。
原文摘要 · Abstract (English)
Recent advancements in reasoning capability of Multimodal Large Language Models (MLLMs) demonstrate its effectiveness in tackling complex visual tasks. However, existing MLLM-based Video Anomaly Detection (VAD) methods remain limited to shallow anomaly descriptions without deep reasoning. In this paper, we propose a new task named Video Anomaly Reasoning (VAR), which aims to enable deep analysis and understanding of anomalies in the video by requiring MLLMs to think explicitly before answering. To this end, we propose Vad-R1, an end-to-end MLLM-based framework for VAR. Specifically, we design a Perception-to-Cognition Chain-of-Thought (P2C-CoT) that simulates the human process of recognizing anomalies, guiding the MLLM to reason anomaly step-by-step. Based on the structured P2C-CoT, we construct Vad-Reasoning, a dedicated dataset for VAR. Furthermore, we propose an improved reinforcement learning algorithm AVA-GRPO, which explicitly incentivizes the anomaly reasoning capability of MLLMs through a self-verification mechanism with limited annotations. Experimental results demonstrate that Vad-R1 achieves superior performance, outperforming both open-source and proprietary models on VAD and VAR tasks. Codes and datasets will be released at https://github.com/wbfwonderful/Vad-R1.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。