首个统一评估场景感知视频异常的基准,提升真实世界异常理解能力
CueBench: Advancing Unified Understanding of Context-Aware Video Anomalies in Real-World
- 构建事件中心的层级分类体系,涵盖14类条件与18类绝对异常事件
- 在174个场景、198个属性上建立统一评测框架,覆盖识别、定位、检测与预测任务
- 提出Cue-R1模型,比现有方法平均提升24%以上,适配多类型视觉语言模型
当前深度模型对真实世界视频异常理解仍停留在表面,缺乏对复杂原则与细微情境的把握。为此,我们提出首个面向场景感知视频异常的统一评估基准CueBench,建立以事件为中心的层级分类体系,涵盖14类条件异常和18类绝对异常事件,基于174个场景与198个属性定义其精细语义。在此基础上,统一评测各类挑战性任务:识别、时序定位、检测与预测。该基准可作为生成-判别及通用-专用视觉语言模型(VLMs)的严格公平评估工具。为应对挑战,我们进一步提出基于R1风格强化微调的Cue-R1模型,采用可验证、任务对齐且层次细化的奖励机制。在CueBench上的实验表明,现有VLMs仍远未达到理想水平,而我们的Cue-R1平均性能超越现有方法超24%。
原文摘要 · Abstract (English)
How far are deep models from real-world video anomaly understanding (VAU)? Current works typically emphasize on detecting unexpected occurrences deviated from normal patterns or comprehending anomalous events with interpretable descriptions. However, they exhibit only a superficial comprehension of real-world anomalies, with limited breadth in complex principles and subtle context that distinguish the anomalies from normalities, e.g., climbing cliffs with safety gear vs. without it. To this end, we introduce CueBench, the first of its kind Benchmark, devoted to Context-aware video anomalies within a Unified Evaluation framework. We comprehensively establish an event-centric hierarchical taxonomy that anchors two core event types: 14 conditional and 18 absolute anomaly events, defined by their refined semantics from diverse contexts across 174 scenes and 198 attributes. Based on this, we propose to unify and benchmark context-aware VAU with various challenging tasks across recognition, temporal grounding, detection, and anticipation. This also serves as a rigorous and fair probing evaluation suite for generative-discriminative as well as generalized-specialized vision-language models (VLMs). To address the challenges underlying CueBench, we further develop Cue-R1 based on R1-style reinforcement fine-tuning with verifiable, task-aligned, and hierarchy-refined rewards in a unified generative manner. Extensive results on CueBench reveal that, existing VLMs are still far from satisfactory real-world anomaly understanding, while our Cue-R1 surpasses these state-of-the-art approaches by over 24% on average.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。