用短窗口滑动学习+LLM自动标注,实现实时暴力检测
Short-Window Sliding Learning for Real-Time Violence Detection via LLM-based Auto-Labeling
- 将视频切为1-2秒片段,用LLM自动生成细粒度标签
- 在RWF-2000上达95.25%准确率,UCF-Crime长视频提升至83.25%
- 适合需要快速响应的智能监控场景
本文提出一种用于实时监控视频中暴力行为检测的短窗口滑动学习框架。与传统长视频训练方法不同,该方法将视频分割为1-2秒的片段,并利用大语言模型(LLM)进行自动标题生成以构建细粒度数据集。每个短片段充分利用所有帧,保持时间连续性,从而精确识别快速发生的暴力事件。实验表明,该方法在RWF-2000数据集上达到95.25%的准确率,并显著提升长视频上的表现(UCF-Crime:83.25%),验证了其在智能监控系统中的强泛化能力和实时适用性。
原文摘要 · Abstract (English)
This paper proposes a Short-Window Sliding Learning framework for real-time violence detection in CCTV footages. Unlike conventional long-video training approaches, the proposed method divides videos into 1-2 second clips and applies Large Language Model (LLM)-based auto-caption labeling to construct fine-grained datasets. Each short clip fully utilizes all frames to preserve temporal continuity, enabling precise recognition of rapid violent events. Experiments demonstrate that the proposed method achieves 95.25\% accuracy on RWF-2000 and significantly improves performance on long videos (UCF-Crime: 83.25\%), confirming its strong generalization and real-time applicability in intelligent surveillance systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。