arXiv:2511.03665cs.CV2025-11被引 3

用事件相机数据实现高效隐私保护的人体动作识别。

A Lightweight 3D-CNN for Event-Based Human Action Recognition with Privacy-Preserving Potential

  • 轻量3D卷积网络建模时空动态,适配边缘部署。
  • 在丰田智能家居与ETRI数据集上达94.17%准确率,F1超0.94。
  • 兼顾隐私保护与性能,适合实际场景的智能监控。

本文提出一种轻量级三维卷积神经网络(3DCNN),用于基于事件视觉数据的人体动作识别(HAR)。传统帧式摄像头会捕捉可识别的个人身份信息,构成隐私隐患;而事件相机仅记录像素亮度变化,具备天然隐私保护特性。所提网络有效建模空间与时间动态,同时保持紧凑结构,适用于边缘设备部署。为缓解类别不平衡并提升泛化能力,采用带类别重加权的焦点损失和针对性数据增强策略。模型在整合丰田智能家居数据集(Toyota Smart Home)与ETRI数据集的复合数据集上训练与评估,实验结果表明,其F1-score达0.9415,整体准确率为94.17%,优于C3D、ResNet3D及MC3_18等基准3D-CNN架构最高达3%。结果验证了事件驱动深度学习在构建高精度、高效率、隐私友好的人体动作识别系统方面的潜力,适用于真实世界边缘应用场景。

原文摘要 · Abstract (English)

This paper presents a lightweight three-dimensional convolutional neural network (3DCNN) for human activity recognition (HAR) using event-based vision data. Privacy preservation is a key challenge in human monitoring systems, as conventional frame-based cameras capture identifiable personal information. In contrast, event cameras record only changes in pixel intensity, providing an inherently privacy-preserving sensing modality. The proposed network effectively models both spatial and temporal dynamics while maintaining a compact design suitable for edge deployment. To address class imbalance and enhance generalization, focal loss with class reweighting and targeted data augmentation strategies are employed. The model is trained and evaluated on a composite dataset derived from the Toyota Smart Home and ETRI datasets. Experimental results demonstrate an F1-score of 0.9415 and an overall accuracy of 94.17%, outperforming benchmark 3D-CNN architectures such as C3D, ResNet3D, and MC3_18 by up to 3%. These results highlight the potential of event-based deep learning for developing accurate, efficient, and privacy-aware human action recognition systems suitable for real-world edge applications.

事件相机动作识别隐私保护轻量模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。