用注意力机制提升低清监控图像中罪犯识别精度
Attention-Guided Efficientnet Architecture For Precise Criminal Identification in Surveillance Images
- 融合CBAM模块与多尺度特征融合,增强劣质图像下的特征学习
- 在SCFace和LFW数据集上达到98.2%识别准确率
- 适合公安监控、安防系统等实际场景应用
由于公共与私人环境中监控摄像头部署日益增多,从监控图像中进行犯罪人员识别已成为智能司法监控系统的关键研究方向。然而,由于图像分辨率低、光照变化、运动模糊、姿态差异、面部遮挡及背景杂乱等因素,基于监控的面部识别仍面临巨大挑战。为此,本文提出一种注意力引导的高效网络(AG-EfficientNet)框架,用于提升监控图像中的精准犯罪识别能力。该框架将EfficientNet-B0与卷积块注意力模块(CBAM)结合,以增强在劣质监控条件下判别性面部特征的学习;引入多尺度监控特征融合策略,保留局部纹理信息与高层语义身份表征;同时采用混合Softmax-Triplet优化机制,提升类间可分性与类内紧凑性,实现鲁棒的犯罪身份区分。实验在Labeled Faces in the Wild(LFW)和SCFace数据集上进行,结果表明,所提框架在识别准确率(98.2%)、精确率(97.9%)、召回率(97.6%)、F1分数(97.7%)和ROC-AUC(0.99)上均优于AlexNet、VGG16、ResNet50、MobileNetV2及标准EfficientNet-B0。Grad-CAM可视化与消融分析进一步验证了注意力引导特征学习策略的有效性。
原文摘要 · Abstract (English)
Criminal identification from surveillance imagery has become a critical research area in intelligent forensic surveillance systems due to the increasing deployment of CCTV cameras in public and private environments. However, surveillance-based face recognition remains highly challenging because of low image resolution, illumination variation, motion blur, pose changes, facial occlusion, and background clutter. To address these limitations, this paper proposes an Attention-Guided EfficientNet (AG-EfficientNet) framework for precise criminal identification in surveillance images. The proposed framework integrates EfficientNet-B0 with Convolutional Block Attention Modules (CBAM) to enhance discriminative facial feature learning under degraded surveillance conditions. In addition, a multi-scale surveillance feature fusion strategy is introduced to preserve both local texture information and high-level semantic identity representations. A hybrid Softmax-Triplet optimization mechanism is further employed to improve inter-class separability and intra-class compactness for robust criminal identity discrimination. The proposed framework was experimentally evaluated using the Labeled Faces in the Wild (LFW) and SCFace datasets. Experimental results demonstrate that the proposed AG-EfficientNet framework achieved superior surveillance recognition performance with an identification accuracy of 98.2%, Precision of 97.9%, Recall of 97.6%, F1-Score of 97.7%, and ROC-AUC of 0.99, outperforming conventional deep learning architectures including AlexNet, VGG16, ResNet50, MobileNetV2, and standard EfficientNet-B0. Furthermore, Grad-CAM visualization and ablation analysis confirm the effectiveness of the proposed attention-guided feature learning strategy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。