用生成模型构建大规模异常视频文本数据集,解决真实数据稀缺与隐私问题。
Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark
- 基于大语言模型生成68类异常事件描述,驱动视频生成模型合成多样视频
- 构建41,315段视频(136万帧)的跨模态数据集,覆盖30类正常行为与68类异常事件
- 可替代真实数据用于安全监控场景的异常检索评估,避免隐私风险
视频异常检索旨在通过自然语言查询定位视频中的异常事件,以支持公共安全。然而现有数据集存在严重局限:(1) 由于现实异常的长尾特性导致数据稀缺;(2) 隐私约束阻碍大规模采集。为一次性解决上述问题,我们提出首个大规模跨模态异常检索基准——SVTA(Synthetic Video-Text Anomaly),利用生成模型克服数据可用性挑战。具体地,我们通过现成大语言模型(LLM)收集并生成涵盖68类异常类别(如投掷、盗窃、射击)的视频描述,这些描述包含常见长尾事件。利用这些文本指导视频生成模型,产出了多样化且高质量的视频。最终,我们的SVTA包含41,315段视频(共136万帧)及其配对字幕,覆盖30类正常活动(如站立、行走、运动)和68类异常事件(如跌倒、斗殴、盗窃、爆炸、自然灾害)。我们采用三种主流视频-文本检索基线全面测试该数据集,揭示其挑战性及评估鲁棒跨模态检索方法的有效性。SVTA在消除真实异常数据采集带来的隐私风险的同时,保持了真实场景的合理性。数据集演示地址:[https://svta-mm.github.io/SVTA.github.io/]。
原文摘要 · Abstract (English)
Video anomaly retrieval aims to localize anomalous events in videos using natural language queries to facilitate public safety. However, existing datasets suffer from severe limitations: (1) data scarcity due to the long-tail nature of real-world anomalies, and (2) privacy constraints that impede large-scale collection. To address the aforementioned issues in one go, we introduce SVTA (Synthetic Video-Text Anomaly benchmark), the first large-scale dataset for cross-modal anomaly retrieval, leveraging generative models to overcome data availability challenges. Specifically, we collect and generate video descriptions via the off-the-shelf LLM (Large Language Model) covering 68 anomaly categories, e.g., throwing, stealing, and shooting. These descriptions encompass common long-tail events. We adopt these texts to guide the video generative model to produce diverse and high-quality videos. Finally, our SVTA involves 41,315 videos (1.36M frames) with paired captions, covering 30 normal activities, e.g., standing, walking, and sports, and 68 anomalous events, e.g., falling, fighting, theft, explosions, and natural disasters. We adopt three widely-used video-text retrieval baselines to comprehensively test our SVTA, revealing SVTA's challenging nature and its effectiveness in evaluating a robust cross-modal retrieval method. SVTA eliminates privacy risks associated with real-world anomaly collection while maintaining realistic scenarios. The dataset demo is available at: [https://svta-mm.github.io/SVTA.github.io/].
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。