动态建模视频异常时间粒度,提升弱监督检测稳定性。
Adaptive Multi-Granularity Temporal Modeling for Weakly Supervised Video Anomaly Detection

- 引入动态位置编码与可学习类别标记,捕捉长程时序依赖
- 通过时序不连续分析自适应分割异常事件,生成判别性表征
- 基于语义相关性的动态融合策略,优于固定Top-k聚合
随着视频监控数据规模远超人工标注能力,弱监督视频异常检测(WSVAD)成为关键研究方向。现有方法多采用多重实例学习(MIL)框架,依赖固定的、手工设计的时序先验进行异常评分,但难以适应真实视频中异常持续时间和动态变化的多样性,常导致片段级预测不稳定。为此,本文提出一种自适应多粒度时序建模框架:首先设计时序精炼模块(TRM),利用动态位置编码和可学习类别标记,建模长程时序依赖并提炼稳定的全局视频表示;其次构建自适应事件分割模块(ESM),通过时序不连续分析识别事件边界,并聚合片段特征生成判别性事件级表示;最后提出自适应相似性融合策略,动态整合片段与事件级异常分数至视频级预测,替代固定的Top-k聚合机制。在两个基准数据集上的大量实验表明,该框架持续优于当前最优方法。
原文摘要 · Abstract (English)
As the scale of video surveillance data outpaces manual annotation capacities, weakly supervised video anomaly detection (WSVAD) has emerged as a critical research frontier. Most existing approaches formulate WSVAD within a Multiple Instance Learning (MIL) framework that relies on rigid, hand-crafted temporal priors to supervise anomaly scoring. However, such formulations exhibit limited adaptability to the wide variation in anomaly durations and temporal dynamics observed in real-world videos, often leading to unstable or unreliable snippet-level predictions. To address this limitation, we propose an adaptive temporal modeling framework for WSVAD that explicitly accounts for variations in video dynamics across multiple temporal granularities. First, we introduce a Temporal Refinement Module (TRM) that leverages dynamic positional encoding and a learnable class token to model long-range temporal dependencies while distilling a stable global video-level representation. Second, to capture anomalous events with varying frequency and duration, we develop an adaptive Event Segmentation Module (ESM) that identifies event boundaries through temporal discontinuity analysis and aggregates snippet features into discriminative event-level representations. Finally, for snippet-level and event-level predictions, we propose an adaptive similarity-based fusion strategy that dynamically integrates anomaly scores into video-level predictions, replacing fixed top-k aggregation heuristics with global semantic relevance. Extensive experiments on two benchmarks demonstrate that the proposed framework consistently outperforms state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。