提出双基准与异常聚焦采样,提升视频异常检测效果。
DUAL-VAD: Dual Benchmarks and Anomaly-Focused Sampling for Video Anomaly Detection
- 用软最大化策略优先采样异常密集片段,兼顾全局覆盖。
- 在UCF-Crime上实现帧级与视频级性能双提升。
- 适合关注异常检测泛化能力的研究者与工程应用。
视频异常检测(VAD)对监控与公共安全至关重要。现有基准仅限于帧级或视频级任务,限制了模型泛化能力的全面评估。本文首次提出基于软最大值的帧分配策略,优先选取异常密集段落,同时保持全视频覆盖,实现时间尺度上的均衡采样。在此基础上构建两个互补基准:基于图像的基准用于评估帧级推理能力,使用代表性帧;基于视频的基准扩展至时序局部片段,并引入异常评分任务。在UCF-Crime数据集上的实验表明,该方法在帧级和视频级均取得显著提升。消融实验证实,异常聚焦采样相比均匀采样和随机采样具有明显优势。
原文摘要 · Abstract (English)
Video Anomaly Detection (VAD) is critical for surveillance and public safety. However, existing benchmarks are limited to either frame-level or video-level tasks, restricting a holistic view of model generalization. This work first introduces a softmax-based frame allocation strategy that prioritizes anomaly-dense segments while maintaining full-video coverage, enabling balanced sampling across temporal scales. Building on this process, we construct two complementary benchmarks. The image-based benchmark evaluates frame-level reasoning with representative frames, while the video-based benchmark extends to temporally localized segments and incorporates an abnormality scoring task. Experiments on UCF-Crime demonstrate improvements at both the frame and video levels, and ablation studies confirm clear advantages of anomaly-focused sampling over uniform and random baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。