用多模态大模型实现伪造人脸的检测、定位与来源追溯
VLForgery Face Triad: Detection, Localization and Attribution via Multimodal Large Language Models

- 引入多模态大模型,实现对扩散模型生成人脸的细粒度分析
- 在真实与伪造图像间识别低级视觉差异,提升分析精度
- 可定位伪造区域并追溯具体生成模型,适合安全与内容审核场景
由扩散模型生成的高质量可控人脸给深度伪造检测带来挑战。现有先进检测器仅提供二元判断,无法实现伪造区域定位、生成方法溯源及成因分析。本文提出基于多模态大模型的细粒度分析框架VLForgery,可实现:1)预测伪造人脸;2)定位部分合成的伪造区域;3)归因具体生成器。为此,我们构建了名为VLF的多样化合成人脸数据集,促进视觉与语言模态交互。提出外源知识引导的描述方法EkCot,利用生成流程知识使MLLM快速捕捉图像内容。设计低级视觉对比管道,识别真实与伪造图像间的固有差异特征,并融入EkCot,形成检测-定位-溯源的结构化分析流程。大量实验表明,VLForgery在检测准确率上优于现有主流方法,且具备区域定位与溯源潜力。
原文摘要 · Abstract (English)
Faces synthesized by diffusion models (DMs) with high-quality and controllable attributes pose a significant challenge for Deepfake detection. Most state-of-the-art detectors only yield a binary decision, incapable of forgery localization, attribution of forgery methods, and providing analysis on the cause of forgeries. In this work, we integrate Multimodal Large Language Models (MLLMs) within DM-based face forensics, and propose a fine-grained analysis triad framework called VLForgery, that can 1) predict falsified facial images; 2) locate the falsified face regions subjected to partial synthesis; and 3) attribute the synthesis with specific generators. To achieve the above goals, we introduce VLF (Visual Language Forensics), a novel and diverse synthesis face dataset designed to facilitate rich interactions between Visual and Language modalities in MLLMs. Additionally, we propose an extrinsic knowledge-guided description method, termed EkCot, which leverages knowledge from the image generation pipeline to enable MLLMs to quickly capture image content. Furthermore, we introduce a low-level vision comparison pipeline designed to identify differential features between real and fake that MLLMs can inherently understand. These features are then incorporated into EkCot, enhancing its ability to analyze forgeries in a structured manner, following the sequence of detection, localization, and attribution. Extensive experiments demonstrate that VLForgery outperforms other state-of-the-art forensic approaches in detection accuracy, with additional potential for falsified region localization and attribution analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。