arXiv:2412.02127cs.CV2024-12

用3D CNN高效检测监控视频中的打斗场景

Streamlining Video Analysis for Efficient Violence Detection

  • 基于X3D模型,结合管状提取与聚类定位打斗片段
  • 在标准数据集上实现高精度打斗分类,有效区分暴力与非暴力事件
  • 适合无人安防、内容审核等实时视频分析场景

本文针对监控摄像头捕获视频中自动化打斗检测的挑战,聚焦于将场景分类为"打斗"或"非打斗"。该任务对提升无人安保系统、在线内容过滤等应用至关重要。我们提出一种基于3D卷积神经网络(3D CNN)的X3D模型方法,结合管状提取、体积裁剪、帧聚合及聚类技术,精确实现打斗场景的定位与分类。大量实验验证了该方法在区分暴力与非暴力事件上的有效性,为构建实用化暴力检测系统提供了重要参考。

原文摘要 · Abstract (English)

This paper addresses the challenge of automated violence detection in video frames captured by surveillance cameras, specifically focusing on classifying scenes as "fight" or "non-fight." This task is critical for enhancing unmanned security systems, online content filtering, and related applications. We propose an approach using a 3D Convolutional Neural Network (3D CNN)-based model named X3D to tackle this problem. Our approach incorporates pre-processing steps such as tube extraction, volume cropping, and frame aggregation, combined with clustering techniques, to accurately localize and classify fight scenes. Extensive experimentation demonstrates the effectiveness of our method in distinguishing violent from non-violent events, providing valuable insights for advancing practical violence detection systems.

视频分析暴力检测3D CNN

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。