arXiv:2508.00312cs.CVcs.AI2025-08被引 2

用生成视频增强弱监督异常检测,低成本提升模型性能

GV-VAD : Exploring Video Generation for Weakly-Supervised Video Anomaly Detection

  • 用文本控制的视频生成模型创造逼真合成视频用于训练
  • 在UCF-Crime数据集上超越现有最先进方法
  • 适合资源有限但需提升异常检测能力的研究者

视频异常检测(VAD)在智能监控等公共安全应用中至关重要。然而,真实世界异常事件稀少、不可预测且标注成本高,导致难以构建大规模数据集,限制了现有模型的性能与泛化能力。为此,我们提出一种生成视频增强的弱监督视频异常检测框架(GV-VAD),利用文本条件视频生成模型生成语义可控且物理合理的合成视频,以低成本扩充训练数据。同时,采用合成样本损失缩放策略,有效控制生成样本对训练的影响。实验表明,该框架在UCF-Crime数据集上优于现有最先进方法。代码已开源:https://github.com/Sumutan/GV-VAD.git。

原文摘要 · Abstract (English)

Video anomaly detection (VAD) plays a critical role in public safety applications such as intelligent surveillance. However, the rarity, unpredictability, and high annotation cost of real-world anomalies make it difficult to scale VAD datasets, which limits the performance and generalization ability of existing models. To address this challenge, we propose a generative video-enhanced weakly-supervised video anomaly detection (GV-VAD) framework that leverages text-conditioned video generation models to produce semantically controllable and physically plausible synthetic videos. These virtual videos are used to augment training data at low cost. In addition, a synthetic sample loss scaling strategy is utilized to control the influence of generated synthetic samples for efficient training. The experiments show that the proposed framework outperforms state-of-the-art methods on UCF-Crime datasets. The code is available at https://github.com/Sumutan/GV-VAD.git.

视频生成异常检测弱监督生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。