arXiv:2512.06845cs.CV2025-12中稿 · ECCV

无需真实异常视频,用生成数据训练出高精度视频异常检测模型

PA-VAD: Diffusion-Based Pseudo-Only Video Anomaly Detection via Domain-Aligned Memory Updates

  • 用扩散模型生成伪异常视频,与真实正常视频联合训练
  • 在ShanghaiTech等数据集上达到98.2%以上AUC,超越现有基线
  • 提出领域对齐机制,解决生成数据带来的特征偏差问题

真实世界中部署视频异常检测(VAD)常受限于异常视频的稀缺、隐私和采集成本。本文提出PA-VAD,一种全新的伪异常仅训练框架,通过将真实正常视频与少量真实正常图像生成的扩散伪异常视频配对,实现无需任何真实异常视频的训练。除了提出生成驱动的训练流程,我们发现伪异常在特征空间存在时空幅度偏差,若不处理会主导多实例学习并损害泛化能力。为此,我们引入域对齐正则化模块(DARM),结合域对齐与使用感知记忆更新,平衡原型覆盖并稳定有偏伪监督下的优化过程。大量实验表明,PA-VAD在ShanghaiTech上达98.2% AUC,UCF-Crime上达82.5%,XD-Violence上达95.1%,且在开放集评估中对未见异常类表现更优。显著优于现有使用真实异常视频的基线,在ShanghaiTech和XD-Violence上分别提升+0.6%和+0.9%,在UCF-Crime上提升+1.9%,证明高质量VAD可不依赖真实异常数据。

原文摘要 · Abstract (English)

Deploying video anomaly detection (VAD) in the real world is often constrained by the scarcity, privacy, and cost of collecting real abnormal footage. We propose PA-VAD, a novel pseudo-only framework that trains an anomaly detector without using any real abnormal videos, by pairing real normal videos with diffusion-synthesized pseudo-abnormal videos generated from a small set of real normal images. Beyond proposing a generation-driven training pipeline, we make a key empirical discovery: pseudo anomalies exhibit a characteristic spatiotemporal magnitude bias in feature space, which can dominate Multiple Instance Learning and degrade generalization if left unaddressed. To counter this pseudo-induced bias, we introduce the Domain-Aligned Regularized Module (DARM), which combines domain alignment with usage-aware memory updates to balance prototype coverage and stabilize optimization under biased pseudo supervision. Extensive experiments demonstrate that PA-VAD achieves 98.2% AUC on ShanghaiTech, 82.5% on UCF-Crime, and 95.1% on XD-Violence, and further improves generalization to unseen anomaly classes in open-set evaluations. Notably, PA-VAD surpasses the best real-abnormal WVAD baselines on ShanghaiTech and XD-Violence by +0.6% and +0.9%, respectively, and improves over the UVAD state of the art on UCF-Crime by +1.9% -showing that high-accuracy VAD is attainable without collecting real abnormal videos.

视频异常检测扩散模型伪异常无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。