arXiv:2511.13276cs.CV2025-11
用视频级别标签训练模型,精准识别监控中的罕见异常行为。
Recognition of Abnormal Events in Surveillance Videos using Weakly Supervised Dual-Encoder Models
- 双编码器融合卷积与变换器特征,通过top-k池化提取关键信息。
- 在UCF-Crime数据集上达到90.7%的AUC,性能显著。
- 适合缺乏细粒度标注的现实监控场景应用。
我们针对仅使用视频级监督信号检测监控视频中罕见且多样的异常事件这一挑战,提出一种双主干框架。该框架通过top-k池化融合卷积神经网络与Transformer表示,有效捕捉时空特征。在UCF-Crime数据集上,模型取得90.7%的受试者工作特征曲线下面积(AUC),验证了方法的有效性。该方法适用于标注成本高的实际监控场景,为弱监督异常检测提供了新思路。
原文摘要 · Abstract (English)
We address the challenge of detecting rare and diverse anomalies in surveillance videos using only video-level supervision. Our dual-backbone framework combines convolutional and transformer representations through top-k pooling, achieving 90.7% area under the curve (AUC) on the UCF-Crime dataset.
异常检测弱监督监控视频双编码器
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。