arXiv:2606.14230cs.CVcs.CL2026-06

融合多域特征提升跨生成器深度伪造检测能力

A Multi-Domain Feature Fusion Framework for Generalizable Deepfake Detection Across Different Generators

论文配图:A Multi-Domain Feature Fusion Framework for Generalizable Deepfake Detection Across Different Generators
图 1 · 摘自论文原文
  • 设计空间-梯度-小波频域联合网络,挖掘多维度伪造痕迹
  • 跨模型检测准确率达79.80%,跨范式达78%,真实数据超75%
  • 适合需要高泛化能力的反深度伪造系统研发人员

深度伪造内容严重威胁隐私、安全与信息真实性。尽管基于空间或频域的方法在GAN生成的深度伪造中表现良好,但对扩散模型生成的内容检测效果不佳。现有方法通常缺乏对多域特征的互补利用,且未系统评估跨生成器鲁棒性。为此,本文提出SGFF-Net(空间-梯度-频率融合网络),在双残差架构中融合空间、梯度与离散小波变换(DWT)频域表示。实验表明,该框架在同数据集上达到98.95%准确率;跨模型与跨范式设置下分别提升至79.80%与78%;结合多源训练与数据增强后,在真实世界数据上从61.50%提升至75.80%。相比单域检测器,该方法通过多域互补特征学习,显著增强跨生成器与跨范式鲁棒性,为构建更可靠的深度伪造检测系统提供实用指导。

原文摘要 · Abstract (English)

Deepfakes are artificially generated images, audio, or videos that threaten privacy, security, and information integrity. Detecting such content is crucial for countering disinformation, as the latest models generate highly realistic content. While spatial- or frequency-based approaches achieve good detection rates on Generative Adversarial Networks (GANs)-based generated deepfakes, they often struggle with recent diffusion model-generated images. In particular, existing approaches rarely exploit complementary multi-domain representations or systematically evaluate cross-generator robustness. To address these challenges, we propose a multi-domain deepfake detection framework called SGFF-Net (Spatial-Gradient-Frequency Fusion Network) that integrates spatial, gradient, and DWT (Discrete Wavelet Transform)-based frequency representations within a dual residual learning architecture. Experimental results show that the SGFF-Net achieves 98.95\% accuracy in intra-dataset evaluation and improves performance in both cross-model (70.46\%) and cross-paradigm (69.94\%) settings. Incorporating multi-source training and data augmentation further enhances robustness, increasing accuracy from 70.46\% to 79.80\% in cross-model evaluation, from 69\% to 78\% in cross-paradigm evaluation, and from 61.50\% to 75.80\% on real-world data. Unlike single-domain detectors, the SGFF-Net learns complementary forensic cues across spatial, gradient, and wavelet-frequency domains, resulting in greater robustness under cross-generator and cross-paradigm evaluation. The results further show that combining multi-domain representations with data diversity and augmentation substantially improves generalization, providing practical insights for developing more reliable deepfake detection systems.

深度伪造检测多域融合扩散模型泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。