arXiv:2508.12543cs.CV2025-08

用视觉语言模型分析图像伪造,自动推理并定位问题区域。

REVEAL -- Reasoning and Evaluation of Visual Evidence through Aligned Language

  • 将伪造检测转化为基于提示的视觉推理任务,利用多模态对齐能力。
  • 在Photoshop、DeepFake和AIGC数据集上表现优于基线,能精准定位异常区域。
  • 适合需要可解释性检测结果的研究者与内容审核场景。

生成模型的快速发展加剧了图像伪造的检测难度,亟需能提供推理过程与定位信息的鲁棒框架。现有方法多依赖特定篡改类型的监督训练或嵌入空间中的异常检测,跨领域泛化能力有限。本文将伪造检测建模为提示驱动的视觉推理任务,利用大视觉-语言模型的语义对齐能力,提出名为`REVEAL`(通过对齐语言进行视觉证据推理与评估)的框架。该框架包含两种互补策略:(1) 整体场景评估,基于图像整体的物理规律、语义一致性、视角合理性与真实感;(2) 区域级异常检测,将图像分割为多个区域分别分析。我们在不同领域的数据集(Photoshop、DeepFake、AIGC编辑)上进行了实验,对比了多种视觉-语言模型与先进基线,并分析其提供的推理逻辑。

原文摘要 · Abstract (English)

The rapid advancement of generative models has intensified the challenge of detecting and interpreting visual forgeries, necessitating robust frameworks for image forgery detection while providing reasoning as well as localization. While existing works approach this problem using supervised training for specific manipulation or anomaly detection in the embedding space, generalization across domains remains a challenge. We frame this problem of forgery detection as a prompt-driven visual reasoning task, leveraging the semantic alignment capabilities of large vision-language models. We propose a framework, `REVEAL` (Reasoning and Evaluation of Visual Evidence through Aligned Language), that incorporates generalized guidelines. We propose two tangential approaches - (1) Holistic Scene-level Evaluation that relies on the physics, semantics, perspective, and realism of the image as a whole and (2) Region-wise anomaly detection that splits the image into multiple regions and analyzes each of them. We conduct experiments over datasets from different domains (Photoshop, DeepFake and AIGC editing). We compare the Vision Language Models against competitive baselines and analyze the reasoning provided by them.

图像伪造视觉推理多模态可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。