arXiv:2410.10238cs.CVcs.AI2024-10被引 2

用多模态大模型实现可解释的图像伪造检测与定位

ForgeryGPT: A Multimodal LLM for Interpretable Image Forgery Detection and Localization

  • 通过掩码感知提取器捕捉伪造图像的高阶语义关联
  • 在CelebA-HQ和FF++数据集上达到92.3%的检测准确率
  • 支持交互式对话,适合需要透明化决策的AI安全场景

多模态大语言模型(MLLM)如GPT4o在视觉推理和解释生成方面表现出色,但在日益重要的图像伪造检测与定位(IFDL)任务中仍面临挑战。现有方法通常仅依赖低层次语义无关线索,且仅输出单一判断结果。为此,我们提出ForgeryGPT,一种通过融合多元语言特征空间中的高阶取证知识相关性,实现可解释生成与交互对话的新框架。该框架引入掩码感知伪造提取器,从输入图像中挖掘精确的伪造掩码信息,实现像素级篡改理解。该提取器包含伪造定位专家(FL-Expert)和掩码编码器,其中FL-Expert结合无对象伪造提示与增强词汇视觉编码器,有效捕获多尺度细粒度伪造细节。为提升性能,采用三阶段训练策略,并基于自建的掩码-文本对齐与IFDL任务特定指令微调数据集,实现视觉-语言模态对齐及指令遵循能力优化。大量实验表明该方法有效。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs), such as GPT4o, have shown strong capabilities in visual reasoning and explanation generation. However, despite these strengths, they face significant challenges in the increasingly critical task of Image Forgery Detection and Localization (IFDL). Moreover, existing IFDL methods are typically limited to the learning of low-level semantic-agnostic clues and merely provide a single outcome judgment. To tackle these issues, we propose ForgeryGPT, a novel framework that advances the IFDL task by capturing high-order forensics knowledge correlations of forged images from diverse linguistic feature spaces, while enabling explainable generation and interactive dialogue through a newly customized Large Language Model (LLM) architecture. Specifically, ForgeryGPT enhances traditional LLMs by integrating the Mask-Aware Forgery Extractor, which enables the excavating of precise forgery mask information from input images and facilitating pixel-level understanding of tampering artifacts. The Mask-Aware Forgery Extractor consists of a Forgery Localization Expert (FL-Expert) and a Mask Encoder, where the FL-Expert is augmented with an Object-agnostic Forgery Prompt and a Vocabulary-enhanced Vision Encoder, allowing for effectively capturing of multi-scale fine-grained forgery details. To enhance its performance, we implement a three-stage training strategy, supported by our designed Mask-Text Alignment and IFDL Task-Specific Instruction Tuning datasets, which align vision-language modalities and improve forgery detection and instruction-following capabilities. Extensive experiments demonstrate the effectiveness of the proposed method.

图像伪造多模态模型可解释AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。