针对女性安全设计的多模态异常检测数据集,填补了现有研究空白。
An Analysis Focused on Womens Safety: Can VAD Models Be Enhanced by a Multi-modal Dataset?

- 构建包含1001段视频的ExtrAnom数据集,覆盖低光、低分辨率等真实场景
- 首次系统标注链抢、跟踪、骚扰等女性相关异常,占比超30%
- 支持跨模态与视觉语言模型验证,适合关注性别敏感性安全研究者
女性安全是现代社会治理的核心议题。当前视频异常检测(VAD)模型常受限于低分辨率监控视频,且现有数据集普遍缺乏对女性相关异常的关注。多数数据集仅包含高亮、高清、近景视频,难以应对链抢、跟踪、不当触碰等隐蔽性犯罪。为此,本文提出新基准ExtrAnom,包含1001段视频(含异常与正常),每段有4类文本标注(1人工+3大模型生成),分为5类犯罪类别。数据涵盖低光(8%)、低分辨率(13%)、远距离拍摄(15%)及白天场景(64%)。异常视频中,跟踪占3.9%,链抢占17.6%,绑架占7.3%,暗杀占2.3%,骚扰占18.9%,正常视频占50%。该数据集支持跨模态与视觉语言模型(VLM)验证。在与主流单模态与多模态数据集(如XD-Violence、UCF-Crime、UCA)及先进方法对比实验中,发现现有数据集无法有效处理女性相关异常。我们认为ExtrAnom可弥补这一关键研究缺口。
原文摘要 · Abstract (English)
Women's safety and security are paramount for a modern society. Often, crimes scenes get recorded through low-resolution CCTV cameras limiting the efficiency of video anomaly detection (VAD) models. Despite substantial progress in VAD research, women-centric anomalies are still underrepresented in datasets as well as in models. Existing datasets primarily cover well-lit, high-resolution and close-shot videos that are inadequate to tackle critical anomalies such as chain snatching, stalking, inappropriate touch, and other subtle forms of crime against women. To address this, we present a new benchmark, referred to as ExtrAnom. It contains 1001 videos (both anomalies and normal) with four textual annotations; one human-generated and three LLM-generated. The videos are arranged in 5 different categories of crimes. The dataset comprises low-light (8%), low-resolution (13%), long-shot (15%), and daytime (64%) anomaly videos. It includes stalking (3.9%), chain snatching (17.6%), kidnapping (7.3%), assassinations (2.3%), harassment (18.9%), and normal (50%) videos. It is possible to perform cross-modal and VLM-based validations using ExtrAnom. We have benchmarked it against popular unimodal and multi-modal VAD datasets (e.g., XD-Violence, UCF-Crime, and UCA) and SOTA methods. Experiments reveal that existing datasets are insufficient to deal with women-centric anomalies. We believe ExtrAnom can fill this critical gap in VAD research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。