为犯罪监控视频分析构建大模型评测基准,推动智能安防发展
A Benchmark for Crime Surveillance Video Analysis with Large Models
- 设计六类问题与多样化问答对,适配大模型开放文本回答
- 覆盖1829段视频,用GPT-4o评估模型表现,验证基准可靠性
- 适合研究视频异常检测与多模态大模型的学者使用
监控视频中的异常行为分析是计算机视觉的关键课题。近年来,多模态大语言模型(MLLMs)在多个领域超越了专用模型。尽管MLLMs具备高度灵活性,但其对异常概念和细节的理解能力仍缺乏充分研究,主要受限于现有基准数据集无法提供符合MLLM特性的问答对及高效评估其开放文本响应的算法。为此,我们提出面向大模型的犯罪监控视频分析基准UCVL,包含1829个视频,并整合了UCF-Crime与UCF-Crime Annotation数据集的标注信息。我们设计了六类问题并生成多样化的问答对,制定详细评估指令,利用OpenAI的GPT-4o进行精确评分。我们对8个主流MLLM(参数量从0.5B到40B)进行了基准测试,结果证明该基准的可靠性。此外,我们在UCVL训练集上微调了LLaVA-OneVision,性能提升验证了数据集在视频异常分析任务中的高质量。
原文摘要 · Abstract (English)
Anomaly analysis in surveillance videos is a crucial topic in computer vision. In recent years, multimodal large language models (MLLMs) have outperformed task-specific models in various domains. Although MLLMs are particularly versatile, their abilities to understand anomalous concepts and details are insufficiently studied because of the outdated benchmarks of this field not providing MLLM-style QAs and efficient algorithms to assess the model's open-ended text responses. To fill this gap, we propose a benchmark for crime surveillance video analysis with large models denoted as UCVL, including 1,829 videos and reorganized annotations from the UCF-Crime and UCF-Crime Annotation datasets. We design six types of questions and generate diverse QA pairs. Then we develop detailed instructions and use OpenAI's GPT-4o for accurate assessment. We benchmark eight prevailing MLLMs ranging from 0.5B to 40B parameters, and the results demonstrate the reliability of this bench. Moreover, we finetune LLaVA-OneVision on UCVL's training set. The improvement validates our data's high quality for video anomaly analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。