构建细粒度视频异常理解新基准,更贴近人类对异常事件的感知。
FineVAU: A Novel Human-Aligned Benchmark for Fine-Grained Video Anomaly Understanding
- 提出三维度异常理解框架:事件、主体与位置,提升描述精度
- 设计新评估指标FVScore,精准衡量视觉关键元素是否被正确识别
- 发布自动构建的细粒度数据集FineW3,适合研究视频理解与人机对齐
视频异常理解(VAU)是一项关注视频中异常事件描述的新任务。现有评估方法多依赖n-gram指标(如BLEU、ROUGE-L)或基于大语言模型的评价,前者无法捕捉自由形式且与视觉相关的回答特征,后者侧重语言质量而非事实相关性,常导致与人类感知不一致的主观判断。本文提出FineVAU,一个聚焦细粒度、领域特定异常理解的新基准。将VAU建模为三重问题:事件(What)、参与主体(Who)和位置(Where)。引入两个核心贡献:一是FVScore,一种新型人类对齐的评估指标,可检测大视觉语言模型输出中关键视觉元素的存在,提供可解释的细粒度反馈;二是FineW3,一个通过结构化全自动流程构建的综合性数据集,扩充了已有标注中的高质量细粒度视觉信息。人工评估显示,本指标在对异常感知的对齐度上优于现有方法。在FineVAU上的实验揭示,尽管大视觉语言模型在粗粒度静态信息和强视觉提示事件上表现良好,但在需要空间与精细时间理解的异常事件上仍存在显著局限。
原文摘要 · Abstract (English)
Video Anomaly Understanding (VAU) is a novel task focused on describing unusual occurrences in videos. Despite growing interest, the evaluation of VAU remains an open challenge. Existing benchmarks rely on n-gram-based metrics (e.g., BLEU, ROUGE-L) or LLM-based evaluation. The first fails to capture the rich, free-form, and visually grounded nature of LVLM responses, while the latter focuses on assessing language quality over factual relevance, often resulting in subjective judgments that are misaligned with human perception. In this work, we address this issue by proposing FineVAU, a new benchmark for VAU that shifts the focus towards rich, fine-grained and domain-specific understanding of anomalous videos. We formulate VAU as a three-fold problem, with the goal of comprehensively understanding key descriptive elements of anomalies in video: events (What), participating entities (Who) and location (Where). Our benchmark introduces a) FVScore, a novel, human-aligned evaluation metric that assesses the presence of critical visual elements in LVLM answers, providing interpretable, fine-grained feedback; and b) FineW3, a novel, comprehensive dataset curated through a structured and fully automatic procedure that augments existing human annotations with high quality, fine-grained visual information. Human evaluation reveals that our proposed metric has a superior alignment with human perception of anomalies in comparison to current approaches. Detailed experiments on FineVAU unveil critical limitations in LVLM's ability to perceive anomalous events that require spatial and fine-grained temporal understanding, despite strong performance on coarse grain, static information, and events with strong visual cues.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。