通过分层融合与多流特征提取,提升对先进生成模型伪造图像的检测能力。
HFMF: Hierarchical Fusion Meets Multi-Stream Models for Deepfake Detection
- 分层跨模态融合视觉变换器与卷积网络特征
- 结合物体级信息与微调卷积网络提升识别精度
- 适用于检测扩散模型等生成的高仿真假图像
深度生成模型的快速发展催生了高度逼真的合成图像,越来越难以与真实数据区分。变分模型、扩散模型和生成对抗网络的广泛应用,使得伪造图像和视频更具欺骗性,严重威胁信息真实性。为此,我们提出HFMF——一种两阶段深度伪造检测框架,结合分层跨模态特征融合与多流特征提取,以增强对前沿生成模型所产图像的检测性能。第一部分通过分层特征融合机制整合视觉变换器与卷积网络;第二部分融合物体级信息与微调后的卷积网络模型输出。最终通过集成深度神经网络融合双路径结果,实现稳健分类。实验表明,该架构在多个数据集上均取得优越性能,并具备良好的校准性与互操作性。
原文摘要 · Abstract (English)
The rapid progress in deep generative models has led to the creation of incredibly realistic synthetic images that are becoming increasingly difficult to distinguish from real-world data. The widespread use of Variational Models, Diffusion Models, and Generative Adversarial Networks has made it easier to generate convincing fake images and videos, which poses significant challenges for detecting and mitigating the spread of misinformation. As a result, developing effective methods for detecting AI-generated fakes has become a pressing concern. In our research, we propose HFMF, a comprehensive two-stage deepfake detection framework that leverages both hierarchical cross-modal feature fusion and multi-stream feature extraction to enhance detection performance against imagery produced by state-of-the-art generative AI models. The first component of our approach integrates vision Transformers and convolutional nets through a hierarchical feature fusion mechanism. The second component of our framework combines object-level information and a fine-tuned convolutional net model. We then fuse the outputs from both components via an ensemble deep neural net, enabling robust classification performances. We demonstrate that our architecture achieves superior performance across diverse dataset benchmarks while maintaining calibration and interoperability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。