融合人脸特征与心率信息,实现轻量级深度伪造统一检测。
A Lightweight and Interpretable Deepfakes Detection Framework
- 结合人脸关键点与新型心率特征,融合提取伪造痕迹。
- 在世界领导人数据集上优于对比方法,准确率达98.7%。
- 模型轻量可解释,适合实际部署与司法场景使用。
近年来,深度伪造的逼真生成与传播对社会生活、公共秩序和法律体系构成严重威胁,包括名人诽谤、选举操纵及作为法庭证据等潜在后果。得益于PyTorch、TensorFlow等开源训练模型、FaceApp和REFACE等视频篡改应用以及低成本计算资源,深度伪造制作门槛大幅降低。现有检测方法多针对特定类型(如换脸、唇形同步、傀儡控制),缺乏统一框架。本文提出一种统一检测框架,融合创新的心率特征与混合人脸关键点特征,更有效地捕捉伪造视频中的面部伪影与原始视频中的自然变化。利用这些特征训练轻量级XGBoost分类器以区分深度伪造与真实视频。在包含三类深度伪造的世界领导人数据集(WLDR)上评估,实验结果表明本框架性能显著优于现有方法。与深度学习模型LSTM-FCN相比,本方法达到相近准确率,但具备更强可解释性。
原文摘要 · Abstract (English)
The recent realistic creation and dissemination of so-called deepfakes poses a serious threat to social life, civil rest, and law. Celebrity defaming, election manipulation, and deepfakes as evidence in court of law are few potential consequences of deepfakes. The availability of open source trained models based on modern frameworks such as PyTorch or TensorFlow, video manipulations Apps such as FaceApp and REFACE, and economical computing infrastructure has easen the creation of deepfakes. Most of the existing detectors focus on detecting either face-swap, lip-sync, or puppet master deepfakes, but a unified framework to detect all three types of deepfakes is hardly explored. This paper presents a unified framework that exploits the power of proposed feature fusion of hybrid facial landmarks and our novel heart rate features for detection of all types of deepfakes. We propose novel heart rate features and fused them with the facial landmark features to better extract the facial artifacts of fake videos and natural variations available in the original videos. We used these features to train a light-weight XGBoost to classify between the deepfake and bonafide videos. We evaluated the performance of our framework on the world leaders dataset (WLDR) that contains all types of deepfakes. Experimental results illustrate that the proposed framework offers superior detection performance over the comparative deepfakes detection methods. Performance comparison of our framework against the LSTM-FCN, a candidate of deep learning model, shows that proposed model achieves similar results, however, it is more interpretable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。