首个异常行为定位基准,评估模型在真实场景中发现罕见事件的能力
UAL-Bench: The First Comprehensive Unusual Activity Localization Benchmark
- 构建包含4个数据集的综合评测平台,覆盖不同异常类型
- VLM-LLM融合模型在短时异常定位与起始时间预测上表现最优
- 提出新指标R@1, TD<=p,更精准衡量异常定位性能
视频中定位异常行为(如人为失误或监控事件)具有实际意义。然而,当前视频理解模型在定位此类事件时表现不佳,可能源于预训练数据中异常样本不足。为探究基础模型在异常行为定位中的能力,我们提出了UAL-Bench——首个综合性异常活动定位基准,包含三个视频数据集:UAG-OOPS、UAG-SSBD、UAG-FunQA,以及一个指令微调数据集OOPS-UAG-Instruct,以提升模型性能。该基准评估三种方法:视频-语言模型(Vid-LLMs)、指令微调的Vid-LLMs,以及视觉-语言模型与大语言模型的新融合方式(VLM-LLM)。结果表明,VLM-LLM在定位短时异常事件及预测其起始时间方面优于Vid-LLMs。我们还提出新指标R@1, TD <= p,以克服现有评估方法的局限性。研究揭示了长时视频(如自闭症诊断场景)带来的挑战,凸显定位技术仍需改进。本工作不仅提供基准,也指明现有模型的关键瓶颈与未来方向。
原文摘要 · Abstract (English)
Localizing unusual activities, such as human errors or surveillance incidents, in videos holds practical significance. However, current video understanding models struggle with localizing these unusual events likely because of their insufficient representation in models' pretraining datasets. To explore foundation models' capability in localizing unusual activity, we introduce UAL-Bench, a comprehensive benchmark for unusual activity localization, featuring three video datasets: UAG-OOPS, UAG-SSBD, UAG-FunQA, and an instruction-tune dataset: OOPS-UAG-Instruct, to improve model capabilities. UAL-Bench evaluates three approaches: Video-Language Models (Vid-LLMs), instruction-tuned Vid-LLMs, and a novel integration of Vision-Language Models and Large Language Models (VLM-LLM). Our results show the VLM-LLM approach excels in localizing short-span unusual events and predicting their onset (start time) more accurately than Vid-LLMs. We also propose a new metric, R@1, TD <= p, to address limitations in existing evaluation methods. Our findings highlight the challenges posed by long-duration videos, particularly in autism diagnosis scenarios, and the need for further advancements in localization techniques. Our work not only provides a benchmark for unusual activity localization but also outlines the key challenges for existing foundation models, suggesting future research directions on this important task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。