arXiv:2507.21507cs.CVcs.MM2025-07AAAI被引 8

首个支持异常定位与理解的视频异常检测基准与框架。

VAGU & GtS: LLM-Based Benchmark and Framework for Joint Video Anomaly Grounding and Understanding

  • 提出GtS框架,通过文本提示分两阶段定位并解析异常。
  • 新基准VAGU包含类别、解释、时间定位和问答标注。
  • 引入联合评估指标JeAUG,兼顾语义与时间精度。

视频异常检测(VAD)旨在识别视频中的异常事件并准确定位其时间区间。现有方法主要分为两类:传统基于DNN的方法侧重时间定位,基于大模型(LLM)的方法强调语义理解。异常理解与定位对全面检测至关重要且可互补,但现有模型与数据集均无法同时支持两项任务。为此,我们提出首个集成两项任务的基准VAGU(Video Anomaly Grounding and Understanding),每个实例包含异常类别、语义解释、精确时间定位及视频问答标注,并提供多选题用于客观评估。基于该数据集,我们提出无需训练的引导式框架Glance then Scrutinize(GtS),先粗粒度定位高概率异常区域,再进行细节解释与边界精修。此外,提出联合评估指标JeAUG,综合评估语义可解释性与时间精度,克服传统指标局限。大量实验验证了本工作在基准、框架与评估指标上的有效性。

原文摘要 · Abstract (English)

Video Anomaly Detection (VAD) aims to identify anomalous events in videos and accurately determine their time intervals. Current VAD methods mainly fall into two categories: traditional DNN-based approaches that focus on temporal localization, and LLM-based approaches that emphasize semantic understanding. Both anomaly understanding and grounding are essential for comprehensive video anomaly detection and can complement each other. However, no existing model or dataset supports both tasks simultaneously. To address this, we introduce VAGU (Video Anomaly Grounding and Understanding), the first benchmark to integrate both tasks. Each VAGU instance includes annotations for anomaly category, semantic explanation, precise temporal grounding and Video QA. We also provide multiple-choice Video QA for objective evaluation. Based on this dataset, we propose Glance then Scrutinize (GtS), a training-free framework guided by textual prompts. The framework first enables coarse localization of high-probability anomalous regions, followed by detailed anomaly interpretation and temporal boundary refinement. Additionally, we propose the JeAUG metric, which jointly evaluates semantic interpretability and temporal precision, overcoming the limitations of traditional metrics. Extensive experiments verify the effectiveness of our benchmark, framework, and evaluation metric.

视频异常大模型基准测试定位理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。