构建大规模细粒度标注的暴力行为检测数据集,提升真实场景下模型泛化能力。
DVD: A Comprehensive Dataset for Advancing Violence Detection in Real-World Scenarios
- 构建500部视频、270万帧的细粒度标注数据集,覆盖多样环境与复杂交互。
- 包含多源摄像头、不同光照条件及丰富元数据,贴近真实暴力事件场景。
- 适合从事视频理解、安全监控与行为识别的研究者使用。
暴力检测(VD)已成为重要研究方向。现有自动化方法受限于数据集数量少、标注粗糙、规模与多样性不足,且缺乏元数据,制约了模型泛化能力。为此,我们提出DVD——一个大规模(500个视频,270万帧)、逐帧标注的暴力检测数据集,涵盖多样环境、多变光照、多源摄像头、复杂社交互动及丰富元数据,旨在捕捉真实世界暴力事件的复杂性。
原文摘要 · Abstract (English)
Violence Detection (VD) has become an increasingly vital area of research. Existing automated VD efforts are hindered by the limited availability of diverse, well-annotated databases. Existing databases suffer from coarse video-level annotations, limited scale and diversity, and lack of metadata, restricting the generalization of models. To address these challenges, we introduce DVD, a large-scale (500 videos, 2.7M frames), frame-level annotated VD database with diverse environments, varying lighting conditions, multiple camera sources, complex social interactions, and rich metadata. DVD is designed to capture the complexities of real-world violent events.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。