arXiv:2503.24267cs.CV2025-03被引 14

FakeScope能精准识别假图并给出可解释的伪造证据分析。

FakeScope: Large Multimodal Expert Model for Transparent AI-Generated Image Forensics

  • 构建多模态推理数据集FakeChain与指令数据集FakeInstruct,训练专家级模型
  • 在闭包与开箱场景下均达顶尖性能,可生成连贯解释与修复建议
  • 零样本下实现定量检测,适用于真实世界未知生成器

生成式AI的迅猛发展带来了内容创作的飞跃,但也催生了高度逼真的虚假图像,严重威胁社会信任。传统检测方法多为二分类,缺乏可解释性。为此,我们提出FakeScope——一个面向AI生成图像取证的大型多模态专家模型。其基础为FakeChain:通过人机协同构建的大规模结构化伪造推理数据集;以及FakeInstruct:迄今最大的多模态指令微调数据集,包含两百万条视觉指令,赋予大模型细腻的取证意识。依托FakeInstruct,FakeScope在闭包与开箱场景中均达到领先性能,不仅能准确识别合成图像,还可提供连贯解释、讨论细粒度伪造痕迹,并提出可操作的优化策略。尤为关键的是,尽管仅用定性标签训练,通过提出的基于令牌的概率估计策略,该模型展现出出色的零样本定量检测能力,对未见生成器具备强泛化性,在真实环境下的表现也稳定可靠。

原文摘要 · Abstract (English)

The rapid and unrestrained advancement of generative artificial intelligence (AI) presents a double-edged sword. While enabling unprecedented creativity, it also facilitates the generation of highly convincing content, undermining societal trust. As image generation techniques become increasingly sophisticated, detecting synthetic images is no longer just a binary task--it necessitates explainable methodologies to enhance trustworthiness and transparency. However, existing detection models primarily focus on classification, offering limited explanatory insights. To address these limitations, we propose FakeScope, an expert large multimodal model (LMM) tailored for AI-generated image forensics, which not only identifies synthetic images with high accuracy but also delivers rich query-contingent forensic insights. At the foundation of our approach is FakeChain, a large-scale dataset containing structured forensic reasoning based on visual trace evidence, constructed via a human-machine collaborative framework. Then we develop FakeInstruct, the largest multimodal instruction tuning dataset to date, comprising two million visual instructions that instill nuanced forensic awareness into LMMs. Empowered by FakeInstruct, FakeScope achieves state-of-the-art performance in both closed-ended and open-ended forensic scenarios. It can accurately distinguish synthetic images, provide coherent explanations, discuss fine-grained forgery artifacts, and suggest actionable enhancement strategies. Notably, despite being trained exclusively on qualitative hard labels, FakeScope demonstrates remarkable zero-shot quantitative detection capability via our proposed token-based probability estimation strategy. Furthermore, it shows robust generalization across unseen image generators and performs reliably under in-the-wild scenarios.

图像取证多模态模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。