用语言引导视觉重建,提升扩散模型伪造人脸的检测与定位能力。
MFVLR: Multi-domain Fine-grained Vision-Language Reconstruction for Generalizable Diffusion Face Forgery Detection and Localization

- 通过语言引导学习细粒度伪造特征,融合多域视觉信息。
- 在跨生成器、跨伪造类型等场景下均超越现有方法。
- 适合关注扩散模型伪造检测的研究者和安全应用开发者。
逼真的面部生成技术快速发展,引发社会与学术界对通用伪造检测与定位方法的迫切需求。以往工作多依赖图像模态捕捉跨域伪造模式,忽视了细粒度文本模态的作用,限制了模型泛化能力;且主要针对GAN生成的图像,难以识别扩散模型合成的伪造人脸。为此,本文提出多域细粒度视觉-语言重建(MFVLR)模型,通过语言引导的伪造表征学习,全面挖掘视觉伪造痕迹,实现对扩散合成人脸的通用检测与定位(DFFDL)。具体地,设计细粒度语言变压器,利用语言重建学习通用细粒度语言嵌入;提出多域视觉编码器,捕获图像与残差域中的通用互补伪造模式;构建视觉解码器以重建图像外观并实现伪造定位;此外,引入创新的即插即用视觉注入模块,增强视觉与语言嵌入的交互。大量实验与可视化结果表明,该模型在跨生成器、跨伪造类型、跨数据集等不同设置下均优于现有最佳方法。
原文摘要 · Abstract (English)
The swift advancement in photo-realistic face generation technology has sparked considerable concerns across society and academia, emphasizing the requirement of generalizable face forgery detection and localization methods. Prior works tend to capture face forgery patterns across multiple domains using image modality, other modalities like fine-grained texts are not comprehensively investigated, which restricts the generalization capability of models. Besides, they usually analyze facial images created by GAN, but struggle to identify and localize those synthesized by diffusion. To solve the problems, in this paper, we devise a novel multi-domain fine-grained vision-language reconstruction (MFVLR) model, which explores comprehensive and diverse visual forgery traces via language-guided face forgery representation learning, to achieve generalizable diffusion-synthesized face forgery detection and localization (DFFDL). Specifically, we devise a fine-grained language transformer that studies general fine-grained language embeddings using language reconstruction. We propose a multi-domain vision encoder to capture general and complementary visual forgery patterns across the image and residual domains. A vision decoder is designed to reconstruct image appearance and achieve forgery localization. Besides, we propose an innovative plug-and-play vision injection module to enhance the interaction between the vision and language embeddings. Extensive experiments and visualizations demonstrate that our network outperforms the state of the art on different settings like cross-generator, cross-forgery, and cross-dataset evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。