arXiv:2410.05900cs.CV2024-10被引 9

通过多时尺度特征学习,提升监控视频异常检测精度。

MTFL: Multi-Timescale Feature Learning for Weakly-Supervised Anomaly Detection in Surveillance Videos

  • 用短、中、长时序片段提取多尺度时空特征
  • 在UCF-Crime上达89.78% AUC,优于现有方法
  • 适合关注复杂动作与长时依赖的异常检测任务

异常事件检测对公共安全至关重要,需融合细粒度运动信息与多时间尺度上下文。为此,本文提出多时尺度特征学习(MTFL)方法,利用视频Swin Transformer从短、中、长时序管段中提取时空特征。实验表明,MTFL在UCF-Crime数据集上达到89.78% AUC,优于当前最优方法;在ShanghaiTech上达95.32% AUC,在XD-Violence上达84.57% AP。此外,我们构建了包含18类、2,591个视频的扩展数据集VADD,覆盖更广泛的现实异常场景。

原文摘要 · Abstract (English)

Detection of anomaly events is relevant for public safety and requires a combination of fine-grained motion information and contextual events at variable time-scales. To this end, we propose a Multi-Timescale Feature Learning (MTFL) method to enhance the representation of anomaly features. Short, medium, and long temporal tubelets are employed to extract spatio-temporal video features using a Video Swin Transformer. Experimental results demonstrate that MTFL outperforms state-of-the-art methods on the UCF-Crime dataset, achieving an anomaly detection performance 89.78% AUC. Moreover, it performs complementary to SotA with 95.32% AUC on the ShanghaiTech and 84.57% AP on the XD-Violence dataset. Furthermore, we generate an extended dataset of the UCF-Crime for development and evaluation on a wider range of anomalies, namely Video Anomaly Detection Dataset (VADD), involving 2,591 videos in 18 classes with extensive coverage of realistic anomalies.

异常检测多尺度特征视频分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。