arXiv:2503.14853cs.CV2025-03ICML被引 30

用大模型理解图像与描述,自动识别并解释深度伪造。

Unlocking the Capabilities of Large Vision-Language Models for Generalizable and Explainable Deepfake Detection

论文配图:Unlocking the Capabilities of Large Vision-Language Models for Generalizable and Explainable Deepfake Detection
图 1 · 摘自论文原文
  • 结合视觉特征与文本描述,分析伪造痕迹的关联性。
  • 在多个数据集上超越现有方法,支持多轮对话交互。
  • 适合需要可解释检测结果的研究与应用者。

当前大型视觉-语言模型(LVLM)在多模态理解方面表现出色,但其在深度伪造检测中的潜力尚未充分挖掘,主要因模型知识与伪造模式不匹配。为此,我们提出一个新框架,旨在释放LVLM在深度伪造检测中的能力。该框架包含知识引导伪造检测器(KFD)、伪造提示学习器(FPL)和大语言模型(LLM)。KFD通过计算图像特征与真实/伪造图像描述嵌入之间的相关性,实现伪造分类与定位。KFD输出经由FPL生成细粒度伪造提示嵌入,再与视觉及问题提示嵌入一起输入至LLM,生成文本化检测结果。在FF++、CDF2、DFD、DFDCP、DFDC和DF40等多个基准上的大量实验表明,本方案在泛化性能上优于当前最先进方法,同时具备多轮对话能力。

原文摘要 · Abstract (English)

Current Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities in understanding multimodal data, but their potential remains underexplored for deepfake detection due to the misalignment of their knowledge and forensics patterns. To this end, we present a novel framework that unlocks LVLMs' potential capabilities for deepfake detection. Our framework includes a Knowledge-guided Forgery Detector (KFD), a Forgery Prompt Learner (FPL), and a Large Language Model (LLM). The KFD is used to calculate correlations between image features and pristine/deepfake image description embeddings, enabling forgery classification and localization. The outputs of the KFD are subsequently processed by the Forgery Prompt Learner to construct fine-grained forgery prompt embeddings. These embeddings, along with visual and question prompt embeddings, are fed into the LLM to generate textual detection responses. Extensive experiments on multiple benchmarks, including FF++, CDF2, DFD, DFDCP, DFDC, and DF40, demonstrate that our scheme surpasses state-of-the-art methods in generalization performance, while also supporting multi-turn dialogue capabilities.

深度伪造检测大模型可解释性多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。