arXiv:2601.10165cs.CV2026-01

构建首个视频异常推理基准,让模型像人一样逐步分析异常事件。

Advancing Adaptive Multi-Stage Video Anomaly Reasoning: A Benchmark Dataset and Method

  • 定义多阶段异常推理新任务,要求模型分步思考视觉感知、因果解释和风险决策。
  • 推出包含8641段视频、超5万样本的大型数据集,标注覆盖不同推理深度。
  • 提出自适应推理模型,支持弱监督下可靠推理,适合智能监控与安全系统研究者。

多模态大语言模型在复杂视频理解任务中展现出潜力,但在视频异常检测与理解(VAD&U)领域,现有方法多局限于异常定位或事后描述,缺乏显式的推理过程、风险意识与决策导向解读。为此,我们提出视频异常推理(VAR)新任务,将异常分析从描述性理解提升为结构化的多阶段推理。该任务要求模型在回答异常相关问题前,依次完成视觉感知、因果解释与风险感知决策。为此,我们构建了一个包含8,641个视频的新数据集,每个视频配有多种问题类型,共超过50,000个样本,是目前规模最大的视频异常推理数据集。标注基于结构化的感知-认知-行动思维链(PerCoAct-CoT),明确领域推理先验。同时,提出异常感知组相对策略优化方法,在弱监督下提升推理可靠性。基于此,我们开发了端到端的多模态大模型方法Vad-R1-Plus,支持自适应层次化推理与风险感知决策。大量实验表明,该基准与方法显著提升了大模型在VAR任务上的推理能力,优于开源与专有基线。

原文摘要 · Abstract (English)

Recent progress in reasoning capabilities of Multimodal Large Language Models(MLLMs) has highlighted their potential for performing complex video understanding tasks. However, in the domain of Video Anomaly Detection and Understanding (VAD&U), existing MLLM-based methods are largely limited to anomaly localization or post-hoc description, lacking explicit reasoning processes, risk awareness, and decision-oriented interpretation. To address this gap, we define a new task termed Video Anomaly Reasoning (VAR), which elevates video anomaly analysis from descriptive understanding to structured, multi-stage reasoning. VAR explicitly requires models to perform progressive reasoning over anomalous events before answering anomaly-related questions, encompassing visual perception, causal interpretation, and risk-aware decision making. To support this task, we present a new dataset with 8,641 videos, where each video is annotated with diverse question types corresponding to different reasoning depths, totaling more than 50,000 samples, making it one of the largest datasets for video anomaly. The annotations are based on a structured Perception-Cognition-Action Chain-of-Thought (PerCoAct-CoT), which formalizes domain-specific reasoning priors for video anomaly understanding. This design enables systematic evaluation of multi-stage and adaptive anomaly reasoning. In addition, we propose Anomaly-Aware Group Relative Policy Optimization to further enhance reasoning reliability under weak supervision. Building upon the proposed task and dataset, we develop an end-to-end MLLM-based VAR model termed Vad-R1-Plus, which supports adaptive hierarchical reasoning and risk-aware decision making. Extensive experiments demonstrate that the proposed benchmark and method effectively advance the reasoning capabilities of MLLMs on VAR tasks, outperforming both open-source and proprietary baselines.

视频异常多模态模型推理机制基准数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。