arXiv:2412.05386cs.CV2024-12中稿 · Signal Image and V…

用人体关键点动态交互建模暴力行为,轻量高效。

DIFEM: Key-points Interaction based Feature Extraction Module for Violence Recognition in Videos

  • 基于人体关键点的动态交互特征提取模块,捕捉速度与关节接近度。
  • 在三个数据集上超越多个现有方法,参数量显著更少。
  • 适合资源受限场景下的实时暴力识别,如监控系统部署。

监控视频中的暴力行为检测对保障公共安全至关重要。现有深度学习方法参数量大,难以在边缘设备部署。本文提出基于人体骨骼关键点的动态交互特征提取模块(DIFEM),通过捕捉特定关节的快速运动及其空间接近性等内在特征,有效建模暴力行为的动态特性。该模块提取速度、关节交叉等特征,结合随机森林、决策树、AdaBoost和k近邻等分类器进行判断。在三个标准暴力识别数据集上进行了充分实验,结果表明该方法在保持较低参数量的前提下,性能优于多个现有最先进方法,具备良好的实用性与部署潜力。

原文摘要 · Abstract (English)

Violence detection in surveillance videos is a critical task for ensuring public safety. As a result, there is increasing need for efficient and lightweight systems for automatic detection of violent behaviours. In this work, we propose an effective method which leverages human skeleton key-points to capture inherent properties of violence, such as rapid movement of specific joints and their close proximity. At the heart of our method is our novel Dynamic Interaction Feature Extraction Module (DIFEM) which captures features such as velocity, and joint intersections, effectively capturing the dynamics of violent behavior. With the features extracted by our DIFEM, we use various classification algorithms such as Random Forest, Decision tree, AdaBoost and k-Nearest Neighbor. Our approach has substantially lesser amount of parameter expense than the existing state-of-the-art (SOTA) methods employing deep learning techniques. We perform extensive experiments on three standard violence recognition datasets, showing promising performance in all three datasets. Our proposed method surpasses several SOTA violence recognition methods.

暴力识别关键点轻量模型动态特征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。