提出新模型,让AI能精准识别未知生成方式的伪造人脸
PVLM: Parsing-Aware Vision Language Model with Dynamic Contrastive Learning for Zero-Shot Deepfake Attribution
- 用面部解析特征捕捉伪造痕迹差异,提升溯源能力
- 在未见过的扩散模型生成伪造脸测试中,准确率超现有方法
- 适合研究深度伪造检测与跨模型溯源的开发者
由于生成模型快速发展,伪造人脸的来源追溯问题受到广泛关注。现有深度伪造溯源(DFA)方法多集中于视觉模态内不同域的交互,而对文本和人脸解析等模态利用不足,且难以细粒度评估对新型生成器(如扩散模型)的泛化性能。本文提出一种基于动态对比学习的解析感知视觉语言模型(PVLM),实现零样本深度伪造溯源(ZSDFA)。我们构建了一个新的细粒度零样本DFA基准,用于评估对未知先进生成器的溯源能力。提出基于视觉-语言模型的PVLM溯源器,利用生成图像中源人脸属性保留程度的差异,设计专门的解析编码器以提取全局面部属性嵌入,通过动态视觉-解析匹配实现解析引导的溯源表示学习。同时引入一种新型对比中心损失,使相关生成器更接近、无关者更远离,增强溯源能力。实验表明,该模型在多种协议评估下均超越当前最优水平。
原文摘要 · Abstract (English)
The challenge of tracing the source attribution of forged faces has gained significant attention due to the rapid advancement of generative models. However, existing deepfake attribution (DFA) works primarily focus on the interaction among various domains in vision modality, and other modalities such as texts and face parsing are not fully explored. Besides, they tend to fail to assess the generalization performance of deepfake attributors to unseen advanced generators like diffusion in a fine-grained manner. In this paper, we propose a novel parsing-aware vision language model with a dynamic contrastive learning (PVLM) method for zero-shot deepfake attribution (ZSDFA), which facilitates effective and fine-grained traceability to unseen advanced generators. Specifically, we conduct a novel and fine-grained ZS-DFA benchmark to evaluate the attribution performance of deepfake attributors to unseen advanced generators like diffusion. Besides, we propose an innovative PVLM attributor based on the vision-language model to capture general and diverse attribution features. We are motivated by the observation that the preservation of source face attributes in facial images generated by GAN and diffusion models varies significantly. We propose to employ the inherent facial attributes preservation differences to capture face parsing-aware forgery representations. Therefore, we devise a novel parsing encoder to focus on global face attribute embeddings, enabling parsing-guided DFA representation learning via dynamic vision-parsing matching. Additionally, we present a novel deepfake attribution contrastive center loss to pull relevant generators closer and push irrelevant ones away, which can be introduced into DFA models to enhance traceability. Experimental results show that our model exceeds the state-of-the-art on the ZS-DFA benchmark via various protocol evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。