arXiv:2503.12381cs.CVcs.MM2025-03

利用耳部生物特征与优化模型提升深度伪造检测精度

Deepfake Detection with Optimized Hybrid Model: EAR Biometric Descriptor via Improved RCNN

  • 结合改进RCNN提取耳部生物特征,构建混合检测模型
  • 在三种数据集上准确率超传统模型,最高达98.7%
  • 适合需要高精度伪造视频识别的安防与媒体场景

深度伪造技术近年广泛用于制造虚假新闻、影视和谣言,通过替换面部信息生成恶意内容。随着AI技术进步,辨别深度伪造图像愈发困难。本文提出一种基于改进RCNN的新型优化混合模型,利用耳部生物特征描述符检测细微耳部运动与形状变化。输入视频经帧提取与预处理(包括缩放、归一化、灰度化及滤波)后,采用Viola-Jones算法进行人脸检测。随后,使用由深度信念网络(DBN)与双向门控循环单元(Bi-GRU)组成的混合模型,基于耳部特征进行检测,并通过优化的分数级融合确定结果。为提升性能,采用自升级水母优化算法(SU-JFO)对模型权重进行最优调节。实验在三个不同数据集上针对压缩、噪声、旋转、姿态和光照四种场景进行评估,结果表明,本方法在准确率、特异性与精确率等指标上均优于传统模型如CNN、SqueezeNet、LeNet、LinkNet、LSTM、DFP [1] 和 ResNext+CNN+LSTM [2],最高准确率达98.7%。

原文摘要 · Abstract (English)

Deepfake is a widely used technology employed in recent years to create pernicious content such as fake news, movies, and rumors by altering and substituting facial information from various sources. Given the ongoing evolution of deepfakes investigation of continuous identification and prevention is crucial. Due to recent technological advancements in AI (Artificial Intelligence) distinguishing deepfakes and artificially altered images has become challenging. This approach introduces the robust detection of subtle ear movements and shape changes to generate ear descriptors. Further, we also propose a novel optimized hybrid deepfake detection model that considers the ear biometric descriptors via enhanced RCNN (Region-Based Convolutional Neural Network). Initially, the input video is converted into frames and preprocessed through resizing, normalization, grayscale conversion, and filtering processes followed by face detection using the Viola-Jones technique. Next, a hybrid model comprising DBN (Deep Belief Network) and Bi-GRU (Bidirectional Gated Recurrent Unit) is utilized for deepfake detection based on ear descriptors. The output from the detection phase is determined through improved score-level fusion. To enhance the performance, the weights of both detection models are optimally tuned using the SU-JFO (Self-Upgraded Jellyfish Optimization method). Experimentation is conducted based on four scenarios: compression, noise, rotation, pose, and illumination on three different datasets. The performance results affirm that our proposed method outperforms traditional models such as CNN (Convolution Neural Network), SqueezeNet, LeNet, LinkNet, LSTM (Long Short-Term Memory), DFP (Deepfake Predictor) [1], and ResNext+CNN+LSTM [2] in terms of various performance metrics viz. accuracy, specificity, and precision.

深度伪造检测生物特征识别混合模型优化算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。