构建58.9万条监控视频问答数据集,推动智能安防理解能力提升
SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models
- 用人工标注+大模型辅助生成,构建大规模监控视频问答数据集
- 涵盖12类认知问题,其中异常检测和因果推理任务准确率不足40%
- 适合研究智能监控、事故分析与安全关键系统的人群使用
监控视频内容理解仍是视觉语言研究中的关键挑战,尤其因其真实场景复杂性、事件动态不规则及安全重要性。本文提出 SurveillanceVQA-589K,目前最大的面向监控领域的开放式视频问答基准,包含589,380对问答,覆盖12种认知多样性问题类型,包括时间推理、因果推断、空间理解与异常解释,涉及正常与异常视频场景。为规模化构建该基准,设计混合标注流程:结合时序对齐的人工撰写描述与基于提示的大视觉语言模型辅助问答生成。同时提出多维度评估协议,用于衡量上下文、时间与因果理解能力。在该框架下评估8个大视觉语言模型,揭示其在因果与异常相关任务上存在显著性能差距,尤其在真实监控场景中表现有限。本基准为智能监控、事件分析与自主决策等安全关键应用提供了实用且全面的研究资源。
原文摘要 · Abstract (English)
Understanding surveillance video content remains a critical yet underexplored challenge in vision-language research, particularly due to its real-world complexity, irregular event dynamics, and safety-critical implications. In this work, we introduce SurveillanceVQA-589K, the largest open-ended video question answering benchmark tailored to the surveillance domain. The dataset comprises 589,380 QA pairs spanning 12 cognitively diverse question types, including temporal reasoning, causal inference, spatial understanding, and anomaly interpretation, across both normal and abnormal video scenarios. To construct the benchmark at scale, we design a hybrid annotation pipeline that combines temporally aligned human-written captions with Large Vision-Language Model-assisted QA generation using prompt-based techniques. We also propose a multi-dimensional evaluation protocol to assess contextual, temporal, and causal comprehension. We evaluate eight LVLMs under this framework, revealing significant performance gaps, especially in causal and anomaly-related tasks, underscoring the limitations of current models in real-world surveillance contexts. Our benchmark provides a practical and comprehensive resource for advancing video-language understanding in safety-critical applications such as intelligent monitoring, incident analysis, and autonomous decision-making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。