arXiv:2501.15656cs.CV2025-01被引 9

用Swin Transformer识别深度伪造图像,准确率达71.29%。

Classifying Deepfakes Using Swin Transformers

  • 采用基于移位窗口的Swin Transformer捕捉图像全局特征
  • 在真实与虚假人脸数据集上达到71.29%测试准确率
  • 验证了Transformer与CNN融合模型的互补优势

深度伪造技术的泛滥对数字媒体的真实性构成严峻挑战,亟需开发鲁棒的检测方法。本研究探索了Swin Transformer——一种利用移位窗口实现自注意力机制的先进架构——在检测与分类深度伪造图像中的应用。基于首尔大学计算智能摄影实验室发布的Real and Fake Face Detection数据集,我们评估了Swin Transformer及其混合模型(如Swin-ResNet、Swin-KNN)识别细微篡改痕迹的能力。结果表明,Swin Transformer优于传统CNN架构(包括VGG16、ResNet18和AlexNet),在测试集上达到71.29%的准确率。此外,研究揭示了混合模型设计的洞见,强调了Transformer与基于CNN方法在深度伪造检测中的互补性。该研究凸显了基于Transformer架构在提升图像篡改检测准确率与泛化能力方面的潜力,为应对深度伪造威胁提供了更有效的对策。

原文摘要 · Abstract (English)

The proliferation of deepfake technology poses significant challenges to the authenticity and trustworthiness of digital media, necessitating the development of robust detection methods. This study explores the application of Swin Transformers, a state-of-the-art architecture leveraging shifted windows for self-attention, in detecting and classifying deepfake images. Using the Real and Fake Face Detection dataset by Yonsei University's Computational Intelligence Photography Lab, we evaluate the Swin Transformer and hybrid models such as Swin-ResNet and Swin-KNN, focusing on their ability to identify subtle manipulation artifacts. Our results demonstrate that the Swin Transformer outperforms conventional CNN-based architectures, including VGG16, ResNet18, and AlexNet, achieving a test accuracy of 71.29%. Additionally, we present insights into hybrid model design, highlighting the complementary strengths of transformer and CNN-based approaches in deepfake detection. This study underscores the potential of transformer-based architectures for improving accuracy and generalizability in image-based manipulation detection, paving the way for more effective countermeasures against deepfake threats.

深度伪造Swin Transformer图像检测视觉对抗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。