arXiv:2510.04630cs.CVcs.AI2025-10被引 2

融合时空频注意力,提升深伪检测泛化能力。

SFANet: Spatial-Frequency Attention Network for Deepfake Detection

  • 分块注意力与频域分割增强关键区域识别
  • 在8个数据集组合上达到最先进准确率
  • 适合需要高鲁棒性的真实场景应用

随着深度伪造技术的兴起,检测篡改媒体已成为紧迫问题。现有方法普遍缺乏跨数据集和生成技术的泛化能力。为此,我们提出一种新型集成框架,结合基于Transformer的架构(如Swin Transformer和ViT)与纹理特征方法,以提升检测精度与鲁棒性。该方法引入创新的数据划分、顺序训练、频域分割、分块注意力及人脸分割技术,有效应对数据不平衡问题,强化眼部、口部等高影响区域特征,并改善泛化性能。在包含8个深度伪造数据集的DFWild-Cup测试集上,模型表现达到当前最优水平。集成结构充分利用了Transformer的全局特征提取能力与纹理方法的可解释性优势。本研究证明,混合模型能有效应对不断演进的深度伪造挑战,为实际应用提供稳健解决方案。

原文摘要 · Abstract (English)

Detecting manipulated media has now become a pressing issue with the recent rise of deepfakes. Most existing approaches fail to generalize across diverse datasets and generation techniques. We thus propose a novel ensemble framework, combining the strengths of transformer-based architectures, such as Swin Transformers and ViTs, and texture-based methods, to achieve better detection accuracy and robustness. Our method introduces innovative data-splitting, sequential training, frequency splitting, patch-based attention, and face segmentation techniques to handle dataset imbalances, enhance high-impact regions (e.g., eyes and mouth), and improve generalization. Our model achieves state-of-the-art performance when tested on the DFWild-Cup dataset, a diverse subset of eight deepfake datasets. The ensemble benefits from the complementarity of these approaches, with transformers excelling in global feature extraction and texturebased methods providing interpretability. This work demonstrates that hybrid models can effectively address the evolving challenges of deepfake detection, offering a robust solution for real-world applications.

深伪检测Transformer注意力机制图像安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。