arXiv:2510.27392cs.CVcs.NE2025-10被引 3

融合伪造痕迹与深度学习,提升深伪检测的鲁棒性与可解释性。

A Hybrid Deep Learning and Forensic Approach for Robust Deepfake Detection

  • 结合噪声残差、压缩痕迹等伪造特征与CNN/ViT深度表征
  • 在多个数据集上F1最高达0.96,压缩下仍保持F1=0.87
  • 热力图重合率达82%,增强检测结果可信度,适合安全应用

生成对抗网络(GANs)和扩散模型的快速发展使合成媒体日益逼真,引发虚假信息、身份欺诈与数字信任危机。现有检测方法或依赖深度学习但泛化能力差、易受干扰,或依赖法医分析但难以应对新型篡改。本文提出一种混合框架,融合噪声残差、JPEG压缩痕迹、频域描述符等法医特征,与卷积神经网络(CNN)及视觉变换器(ViT)的深度表示。在FaceForensics++、Celeb-DF v2、DFDC基准数据集上,模型表现优于单一方法与现有最先进混合方法,分别取得F1分数0.96、0.82、0.77。鲁棒性测试显示,在压缩(QF=50时F1=0.87)、对抗扰动(AUC=0.84)及未见篡改场景(F1=0.79)下性能稳定。可解释性分析表明,Grad-CAM与法医热力图在82%案例中与真实篡改区域重合,提升透明度与用户信任。结果验证了混合方法在适应性与可解释性间的平衡优势,为构建稳健可信的深伪检测系统提供有效路径。

原文摘要 · Abstract (English)

The rapid evolution of generative adversarial networks (GANs) and diffusion models has made synthetic media increasingly realistic, raising societal concerns around misinformation, identity fraud, and digital trust. Existing deepfake detection methods either rely on deep learning, which suffers from poor generalization and vulnerability to distortions, or forensic analysis, which is interpretable but limited against new manipulation techniques. This study proposes a hybrid framework that fuses forensic features, including noise residuals, JPEG compression traces, and frequency-domain descriptors, with deep learning representations from convolutional neural networks (CNNs) and vision transformers (ViTs). Evaluated on benchmark datasets (FaceForensics++, Celeb-DF v2, DFDC), the proposed model consistently outperformed single-method baselines and demonstrated superior performance compared to existing state-of-the-art hybrid approaches, achieving F1-scores of 0.96, 0.82, and 0.77, respectively. Robustness tests demonstrated stable performance under compression (F1 = 0.87 at QF = 50), adversarial perturbations (AUC = 0.84), and unseen manipulations (F1 = 0.79). Importantly, explainability analysis showed that Grad-CAM and forensic heatmaps overlapped with ground-truth manipulated regions in 82 percent of cases, enhancing transparency and user trust. These findings confirm that hybrid approaches provide a balanced solution, combining the adaptability of deep models with the interpretability of forensic cues, to develop resilient and trustworthy deepfake detection systems.

深伪检测混合模型可解释性法医分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。