用频域掩码提升检测器泛化能力,兼顾效率与可持续性
Towards Sustainable Universal Deepfake Detection with Frequency-Domain Masking
- 在频域进行随机掩码和几何变换,弱化对空间特征的依赖
- 在多种生成模型上实现顶尖泛化性能,剪枝后仍保持稳定
- 适合需要高效、可扩展检测方案的研究者和应用方
通用深度伪造检测旨在识别多种生成模型(包括未见过的)生成的图像。这要求检测器具备对新出现深度伪造的强大泛化能力,同时降低计算开销以支持大规模筛查,契合绿色AI目标。本文探索频域掩码作为训练策略。与依赖空间特征或大规模预训练模型的传统方法不同,本方法引入随机掩码和几何变换,重点利用频域掩码优异的泛化特性。实验表明,频域掩码不仅提升了对多样化生成器的检测准确率,且在显著模型剪枝下仍保持性能,提供一种可扩展、资源友好的解决方案。该方法在GAN和扩散模型生成的图像数据集上达到当前最优泛化效果,并在结构化剪枝下表现出持续鲁棒性。结果表明,基于频域的掩码是迈向可持续、强泛化深度伪造检测的重要一步。代码与模型已开源。
原文摘要 · Abstract (English)
Universal deepfake detection aims to identify AI-generated images across a broad range of generative models, including unseen ones. This requires robust generalization to new and unseen deepfakes, which emerge frequently, while minimizing computational overhead to enable large-scale deepfake screening, a critical objective in the era of Green AI. In this work, we explore frequency-domain masking as a training strategy for deepfake detectors. Unlike traditional methods that rely heavily on spatial features or large-scale pretrained models, our approach introduces random masking and geometric transformations, with a focus on frequency masking due to its superior generalization properties. We demonstrate that frequency masking not only enhances detection accuracy across diverse generators but also maintains performance under significant model pruning, offering a scalable and resource-conscious solution. Our method achieves state-of-the-art generalization on GAN- and diffusion-generated image datasets and exhibits consistent robustness under structured pruning. These results highlight the potential of frequency-based masking as a practical step toward sustainable and generalizable deepfake detection. Code and models are available at https://github.com/chandlerbing65nm/FakeImageDetection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。