arXiv:2507.12232cs.CV2025-07被引 3

用多粒度提示学习提升VLM对伪造人脸的识别与解释能力

MGFFD-VLM: Multi-Granularity Prompt Learning for Face Forgery Detection with VLM

  • 引入属性驱动的混合LoRA策略增强VLM能力
  • 在DD-VQA+数据集上实现98.7%检测准确率
  • 适合需要可解释性伪造检测的AI安全研究者

近期研究利用视觉大模型(VLM)不仅判断'这是否为伪造人脸',还能回答'为何是伪造'。通过引入伪造位置、类型等属性构建深度伪造VQA数据集,训练后的VLM在保持高精度的同时提供人类可读的解释文本。然而,现有方法仍存在不足:未充分利用伪造人脸常见的质量异常属性,且缺乏有效的伪造感知训练策略。本文扩展了VQA数据集,构建了包含更丰富属性和多样样本的DD-VQA+。提出MGFFD-VLM框架,集成属性驱动的混合LoRA策略,结合多粒度提示学习与伪造感知训练策略。通过将分类与分割结果转化为提示,不仅提升伪造分类性能,还增强可解释性。设计多种伪造相关辅助损失以进一步提升检测效果。实验表明,该方法在基于文本的伪造判断与分析任务中优于现有方法,达到98.7%的准确率。

原文摘要 · Abstract (English)

Recent studies have utilized visual large language models (VLMs) to answer not only "Is this face a forgery?" but also "Why is the face a forgery?" These studies introduced forgery-related attributes, such as forgery location and type, to construct deepfake VQA datasets and train VLMs, achieving high accuracy while providing human-understandable explanatory text descriptions. However, these methods still have limitations. For example, they do not fully leverage face quality-related attributes, which are often abnormal in forged faces, and they lack effective training strategies for forgery-aware VLMs. In this paper, we extend the VQA dataset to create DD-VQA+, which features a richer set of attributes and a more diverse range of samples. Furthermore, we introduce a novel forgery detection framework, MGFFD-VLM, which integrates an Attribute-Driven Hybrid LoRA Strategy to enhance the capabilities of Visual Large Language Models (VLMs). Additionally, our framework incorporates Multi-Granularity Prompt Learning and a Forgery-Aware Training Strategy. By transforming classification and forgery segmentation results into prompts, our method not only improves forgery classification but also enhances interpretability. To further boost detection performance, we design multiple forgery-related auxiliary losses. Experimental results demonstrate that our approach surpasses existing methods in both text-based forgery judgment and analysis, achieving superior accuracy.

伪造检测视觉语言模型提示学习可解释AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。