arXiv:2412.07183cs.CVcs.AI2024-12IJCV被引 9

构建视频异常因果理解新基准,回答‘发生了什么、为何发生、严重程度如何’。

Exploring What Why and How: A Multifaceted Benchmark for Causation Understanding of Video Anomaly

  • 设计三重标注体系,覆盖异常类型、成因解释与影响描述。
  • 提出硬提示与软提示结合的提示工程方法,提升模型对时空关联的捕捉能力。
  • 开发专用评估指标AnomEval,更贴近人类判断标准。

视频异常理解(VAU)的最新进展推动了交通监控、工业自动化等领域的应用。现有基准多聚焦异常检测与定位,而本研究深入探索实际应用中的因果理解问题,提出涵盖‘发生了什么’‘为何发生’‘严重程度如何’三个维度的视频异常因果理解基准ECVA。每个视频均配有三类人工标注:异常类型、起止时间与事件描述;自然语言形式的成因解释;以及反映异常影响的自由文本。基于此,我们提出一种新型提示驱动方法,采用‘硬提示’引导模型关注异常片段关键区域,‘软提示’建模片段内时空关系。同时设计专用评估指标AnomEval,契合人类判断标准,全面评估各类视频大模型在因果理解任务上的表现。实验验证了该方法的有效性,并指明未来研究方向。

原文摘要 · Abstract (English)

Recent advancements in video anomaly understanding (VAU) have opened the door to groundbreaking applications in various fields, such as traffic monitoring and industrial automation. While the current benchmarks in VAU predominantly emphasize the detection and localization of anomalies. Here, we endeavor to delve deeper into the practical aspects of VAU by addressing the essential questions: "what anomaly occurred?", "why did it happen?", and "how severe is this abnormal event?". In pursuit of these answers, we introduce a comprehensive benchmark for Exploring the Causation of Video Anomalies (ECVA). Our benchmark is meticulously designed, with each video accompanied by detailed human annotations. Specifically, each instance of our ECVA involves three sets of human annotations to indicate "what", "why" and "how" of an anomaly, including 1) anomaly type, start and end times, and event descriptions, 2) natural language explanations for the cause of an anomaly, and 3) free text reflecting the effect of the abnormality. Building upon this foundation, we propose a novel prompt-based methodology that serves as a baseline for tackling the intricate challenges posed by ECVA. We utilize "hard prompt" to guide the model to focus on the critical parts related to video anomaly segments, and "soft prompt" to establish temporal and spatial relationships within these anomaly segments. Furthermore, we propose AnomEval, a specialized evaluation metric crafted to align closely with human judgment criteria for ECVA. This metric leverages the unique features of the ECVA dataset to provide a more comprehensive and reliable assessment of various video large language models. We demonstrate the efficacy of our approach through rigorous experimental analysis and delineate possible avenues for further investigation into the comprehension of video anomaly causation.

视频异常因果理解提示工程评估指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。