用多模态细粒度特征提升扩散人脸伪造检测泛化能力
MFCLIP: Multi-modal Fine-grained CLIP for Generalizable Diffusion Face Forgery Detection
- 融合图像与噪声模态,通过语言引导学习细粒度伪造特征
- 在跨生成器、跨数据集等场景下显著超越现有方法
- 适合关注伪造检测泛化性与多模态建模的研究者
逼真的面部生成技术快速发展,引发社会与学术界对人脸伪造检测(FFD)的迫切需求。现有方法主要依赖图像模态,忽视细粒度噪声与文本信息,导致泛化能力受限;且多数模型难以识别未见过的扩散模型生成伪造图像。为此,本文提出多模态细粒度CLIP(MFCLIP)模型,利用对比语言-图像预训练(CLIP)实现通用扩散人脸伪造检测(DFFD)。设计细粒度语言编码器(FLE),从分层文本提示中提取全局语言特征;构建多模态视觉编码器(MVE),捕获全局图像伪造嵌入与来自最丰富补丁的细粒度噪声模式,并融合以挖掘通用视觉伪造痕迹。此外,提出即插即用的样本对注意力(SPA)机制,强化相关负样本对,抑制无关对,实现跨模态样本对更灵活对齐。大量实验与可视化表明,该模型在跨生成器、跨伪造类型、跨数据集评估中均优于现有方法。
原文摘要 · Abstract (English)
The rapid development of photo-realistic face generation methods has raised significant concerns in society and academia, highlighting the urgent need for robust and generalizable face forgery detection (FFD) techniques. Although existing approaches mainly capture face forgery patterns using image modality, other modalities like fine-grained noises and texts are not fully explored, which limits the generalization capability of the model. In addition, most FFD methods tend to identify facial images generated by GAN, but struggle to detect unseen diffusion-synthesized ones. To address the limitations, we aim to leverage the cutting-edge foundation model, contrastive language-image pre-training (CLIP), to achieve generalizable diffusion face forgery detection (DFFD). In this paper, we propose a novel multi-modal fine-grained CLIP (MFCLIP) model, which mines comprehensive and fine-grained forgery traces across image-noise modalities via language-guided face forgery representation learning, to facilitate the advancement of DFFD. Specifically, we devise a fine-grained language encoder (FLE) that extracts fine global language features from hierarchical text prompts. We design a multi-modal vision encoder (MVE) to capture global image forgery embeddings as well as fine-grained noise forgery patterns extracted from the richest patch, and integrate them to mine general visual forgery traces. Moreover, we build an innovative plug-and-play sample pair attention (SPA) method to emphasize relevant negative pairs and suppress irrelevant ones, allowing cross-modality sample pairs to conduct more flexible alignment. Extensive experiments and visualizations show that our model outperforms the state of the arts on different settings like cross-generator, cross-forgery, and cross-dataset evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。