融合空间与频域特征,用多阶段训练提升深度伪造检测准确率。
CAE-Net: Generalized Deepfake Image Detection using Convolution and Attention Mechanisms with Spatial and Frequency Domain Features
- 结合卷积与注意力机制,融合三种主干网络和小波特征
- 在5:1不均衡数据集上达94.46%准确率和97.60% AUC
- 多阶段非重叠子集训练增强泛化性,适合对抗攻击场景
深度伪造的传播带来严重安全威胁,亟需可靠的检测方法。然而多样化的生成技术与数据集中的类别不平衡带来了挑战。本文提出CAE-Net,一种基于卷积与注意力机制的加权集成网络,融合空间域与频域特征以实现高效深度伪造检测。该架构结合EfficientNet、Data-Efficient Image Transformer (DeiT) 和 ConvNeXt,并引入小波特征以学习互补表征。我们在具有5:1假/真类不平衡的IEEE Signal Processing Cup 2025(DF-Wild Cup)数据集上评估了CAE-Net。为应对不平衡问题,我们设计了一种多阶段非重叠子集训练策略,在不重叠的假图像子集上分阶段训练,同时保留各阶段知识。所提方法取得94.46%准确率和97.60% AUC,优于传统类别平衡方法。可视化表明网络聚焦于有意义的面部区域,且集成设计对对抗攻击表现出鲁棒性,使CAE-Net成为可靠且通用的深度伪造检测框架。
原文摘要 · Abstract (English)
The spread of deepfakes poses significant security concerns, demanding reliable detection methods. However, diverse generation techniques and class imbalance in datasets create challenges. We propose CAE-Net, a Convolution- and Attention-based weighted Ensemble network combining spatial and frequency-domain features for effective deepfake detection. The architecture integrates EfficientNet, Data-Efficient Image Transformer (DeiT), and ConvNeXt with wavelet features to learn complementary representations. We evaluated CAE-Net on the diverse IEEE Signal Processing Cup 2025 (DF-Wild Cup) dataset, which has a 5:1 fake-to-real class imbalance. To address this, we introduce a multistage disjoint-subset training strategy, sequentially training the model on non-overlapping subsets of the fake class while retaining knowledge across stages. Our approach achieved $94.46\%$ accuracy and a $97.60\%$ AUC, outperforming conventional class-balancing methods. Visualizations confirm the network focuses on meaningful facial regions, and our ensemble design demonstrates robustness against adversarial attacks, positioning CAE-Net as a dependable and generalized deepfake detection framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。