通过融合表情与语音特征,实现高效实时的深度伪造检测
ExpSpeech-Net: Multimodal Fusion of Expression and Speech for Deepfake Detection

- 结合表情与语音双模态信息,用轻量级网络捕捉深层伪造痕迹
- 在公开数据集上达到94.5%准确率,99.3%精确率,优于传统方法
- 适合部署于移动端或实时系统,适用于日常内容审核场景
深度伪造视频正严重威胁在线内容可信度。现有检测方法多依赖复杂高耗能模型,限制实际应用。本文提出ExpSpeech-Net深伪检测框架(SqN-R-DFD),采用SqueezeNet与循环神经网络(RNN)作为主干网络,同时分析面部表情与语音模式,实现轻量化高效检测。该方法引入ISLBT图像特征与MPNCC信号特征,并结合沙燕辅助粘菌算法(SASMA)进行智能特征选择,确保输入数据最优平衡。通过融合双模态信息,有效捕捉深伪视频中的细微不一致。实验表明,该模型在测试集上实现94.5%准确率、99.3%精确率与96.8%F-measure,显著优于传统方法,验证了多模态融合结合智能预处理与特征选择在实际、实时深伪检测中的可行性。
原文摘要 · Abstract (English)
Deepfake videos are increasingly challenging the credibility of online content. Many existing detection methodology relies on complex, resource-intensive models, which limit their practical use. The study introduces the ExpSpeech-Net deepfake detection (SqN-R-DFD) model, which utilizes SqueezeNet and RNN (Recurrent Neural Network) as its backbone, providing a lightweight and efficient deepfake detection framework that simultaneously analyzes facial expressions and speech patterns. The approach incorporates advanced feature extraction, such as ISLBT-based features for image and MPNCC for signals, along with a smart feature-selection strategy using SASMA (Sandpiper-Assisted Slime Mould Algorithm), ensuring optimal and balanced input to the detection models. By combining SqueezeNet and an RNN, subtle inconsistencies in deepfake videos are captured effectively. The framework achieves 94.5% accuracy, precision of 99.3%, and F-measure of 96.8%, outperforming conventional methods. This demonstrates that integrating multiple modalities with intelligent preprocessing and feature selection enables practical, real-time deepfake detection suitable for everyday applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。