arXiv:2508.01402cs.CV2025-08被引 6

用多模态大模型实现可解释的AI生成图像检测,让机器像人一样分析伪造痕迹。

ForenX: Towards Explainable AI-Generated Image Detection with Multimodal Large Language Models

  • 借助多模态大模型,结合专用取证提示,聚焦伪造特征。
  • 在两个主流基准上表现优异,解释质量获主观评估验证。
  • 适合需要可信溯源的图像审核、媒体鉴真等场景。

生成模型的进步使得AI生成图像与真实图像在视觉上几乎无法区分。尽管已有大量研究通过分类器检测此类图像,但现有方法与人类认知式鉴定之间仍存在差距。我们提出ForenX,一种不仅能判断图像真伪,还能提供符合人类思维逻辑解释的新方法。ForenX利用强大的多模态大语言模型(MLLMs)分析并解读取证线索,并通过引入专门设计的取证提示,引导模型关注伪造相关属性,克服了标准MLLM在伪造检测中的局限性。该方法不仅提升了检测泛化能力,还使模型能生成准确、相关且全面的解释。此外,我们构建了ForgeryReason数据集,通过大模型代理与人工标注团队协作标注AI生成图像中的伪造证据描述,获得高质量训练数据,验证了少量人工标注即可显著提升解释质量。我们在两个主要基准上评估了ForenX的有效性,其可解释性通过全面的主观评价得到验证。

原文摘要 · Abstract (English)

Advances in generative models have led to AI-generated images visually indistinguishable from authentic ones. Despite numerous studies on detecting AI-generated images with classifiers, a gap persists between such methods and human cognitive forensic analysis. We present ForenX, a novel method that not only identifies the authenticity of images but also provides explanations that resonate with human thoughts. ForenX employs the powerful multimodal large language models (MLLMs) to analyze and interpret forensic cues. Furthermore, we overcome the limitations of standard MLLMs in detecting forgeries by incorporating a specialized forensic prompt that directs the MLLMs attention to forgery-indicative attributes. This approach not only enhance the generalization of forgery detection but also empowers the MLLMs to provide explanations that are accurate, relevant, and comprehensive. Additionally, we introduce ForgReason, a dataset dedicated to descriptions of forgery evidences in AI-generated images. Curated through collaboration between an LLM-based agent and a team of human annotators, this process provides refined data that further enhances our model's performance. We demonstrate that even limited manual annotations significantly improve explanation quality. We evaluate the effectiveness of ForenX on two major benchmarks. The model's explainability is verified by comprehensive subjective evaluations.

AI检测可解释性多模态图像取证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。