arXiv:2409.18301cs.CVcs.AI2024-09被引 27

用小波变换增强视觉模型,提升对未知深伪人脸的检测能力

Wavelet-Driven Generalizable Framework for Deepfake Face Forgery Detection

  • 结合小波变换与CLIP预训练视觉模型,捕捉图像空间与频率特征
  • 跨数据集检测平均AUC达0.749,对未知生成模型深伪检测AUC达0.893
  • 适合关注通用性深伪检测的算法研究者和安全应用开发者

随着深度生成模型的发展,数字图像伪造技术日益复杂,现有深伪检测方法在面对来源不明的伪造图像时面临挑战。为应对这一问题,我们提出Wavelet-CLIP框架,将小波变换与基于CLIP预训练的ViT-L/14视觉特征相结合。该方法通过小波变换深入分析图像的空间与频率特征,显著提升模型对复杂深伪的识别能力。我们在多个标准扩散模型生成的未见图像上进行了全面评估,结果显示,该方法在跨数据集泛化任务中平均AUC达到0.749,在对抗未知深伪的鲁棒性测试中AUC高达0.893,优于所有对比方法。代码已开源:https://github.com/lalithbharadwajbaru/Wavelet-CLIP。

原文摘要 · Abstract (English)

The evolution of digital image manipulation, particularly with the advancement of deep generative models, significantly challenges existing deepfake detection methods, especially when the origin of the deepfake is obscure. To tackle the increasing complexity of these forgeries, we propose \textbf{Wavelet-CLIP}, a deepfake detection framework that integrates wavelet transforms with features derived from the ViT-L/14 architecture, pre-trained in the CLIP fashion. Wavelet-CLIP utilizes Wavelet Transforms to deeply analyze both spatial and frequency features from images, thus enhancing the model's capability to detect sophisticated deepfakes. To verify the effectiveness of our approach, we conducted extensive evaluations against existing state-of-the-art methods for cross-dataset generalization and detection of unseen images generated by standard diffusion models. Our method showcases outstanding performance, achieving an average AUC of 0.749 for cross-data generalization and 0.893 for robustness against unseen deepfakes, outperforming all compared methods. The code can be reproduced from the repo: \url{https://github.com/lalithbharadwajbaru/Wavelet-CLIP}

深伪检测小波变换视觉模型泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。