构建视频事件检测新方法的三支柱框架,助力研究定位与评估。
Event Detection in Videos: A Framework for the Development of New Methods

- 以数据集、评估方法、部署场景为三大支柱构建开发框架
- 强调缺乏大规模数据集和严谨评估导致方法比较失真
- 适合从事视频分析、智能监控的研究者参考
视频中的事件检测是视频监控的核心任务,旨在像素级、帧级或片段级检测事件。针对不同环境、应用场景及采集技术,已有大量检测方法被提出。尽管尝试对这些算法按性能或实时性进行分类,但由于缺乏大规模数据集和严格的评估方法,此类比较存在偏差,影响了方法的进一步发展。鉴于现有方法的多样性,我们主张研究者需在这一丰富背景下明确自身工作的定位。为此,本文提出一个严谨的视频事件检测新方法开发框架,包含三大核心支柱:数据集、性能评估和方法部署场景。
原文摘要 · Abstract (English)
Event detection tasks in videos, the most important aspect of video surveillance, aim to detect events either at the pixel-level, frame-level, or clip-level. Plenty of methods intended for event detection in different environments, for various applications, and within different acquisition techniques were introduced. Naturally, the attempts were made as well to classify these algorithms in terms of detection of performance or in terms of real-time abilities. Nevertheless, the lack of a large-scale dataset as well as rigorous performance evaluation methods have biased such comparisons as well as the development of the methods. Given the diversity of existing approaches, we believe it is essential for researchers to position their work within such a rich landscape. Thus, we propose a rigorous framework for developing new methods in event detection for videos. Specifically, this framework is based on three main pillars: datasets, performance evaluation, and scenarios for deploying methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。