arXiv:2608.16259cs.CVcs.AI2026-08中稿 · ACMMM 2026

让AI生成图检测结果可验证,从猜想到有据可查。

Defake-o3: From Speculative Rationales to Verifiable Evidence for Explainable AIGI Detection

论文配图:Defake-o3: From Speculative Rationales to Verifiable Evidence for Explainable AIGI Detection
图 1 · 摘自论文原文
  • 通过视觉搜索+证据校验器,逐级聚焦可疑区域
  • 在多个数据集上准确率提升,解释更局部、可验证
  • 适合需要可信解释的图像审核与内容安全场景

图像生成模型快速发展,亟需既准确又可解释、可靠的AIGI检测方法。现有基于多模态大模型的检测器虽能生成自然语言解释,但常产生推测性理由:依赖模糊或虚构的伪造痕迹,忽略最新生成器的细微局部缺陷,且无法提供可视化的可验证证据。本文提出Defake-o3,一种从推测性解释转向可验证证据的可解释检测方法。它结合交互式视觉搜索与验证器引导的证据对齐机制:模型迭代地放大可疑区域以检查细粒度细节,同时一个由人类验证标注训练的证据验证器,通过强化学习奖励机制,鼓励真实证据并惩罚无根据声明。为支持该目标,我们构建了GroundFake数据集,包含局部边界框证据、基于视觉定位和瑕疵特异性的真人验证、修正的推理轨迹以及有效/无效证据监督。此外,我们引入FakeFrontier,一个基于真实图像和10个最新生成器输出的分布外基准,配合基于MLLM的协议来评估证据质量和说服力。在GroundFake、FakeFrontier及额外分布外基准上的实验表明,Defake-o3在检测准确率和解释质量上均有提升,生成的证据更具局部性、可验证性和说服力。

原文摘要 · Abstract (English)

The rapid progress of image generation models calls for AI-generated image (AIGI) detectors that are not only accurate but also explainable and reliable. While MLLM-based detectors can provide natural language explanations, existing methods often generate speculative rationales: they rely on vague or hallucinated artifacts, miss subtle localized flaws from the latest generators, and fail to provide evidence that can be visually verified. We present Defake-o3, an explainable AIGI detector that moves from speculative rationales to verifiable evidence. It combines interactive visual search with verifier-guided evidence alignment: the model iteratively zooms into suspicious regions to inspect fine-grained details, while an Evidence Verifier, trained from human verification annotations, provides reinforcement learning rewards that favor grounded evidence and penalize baseless claims. To support this objective, we construct GroundFake, a dataset designed for grounded explainable detection, with localized bounding-box evidence, human verification based on visual grounding and artifact specificity, corrected reasoning trajectories, and valid/invalid evidence supervision. We further introduce FakeFrontier, an out-of-distribution benchmark built from real images and outputs of 10 recent generators, together with an MLLM-based protocol for evaluating evidence quality and persuasiveness. Experiments on GroundFake, FakeFrontier, and additional out-of-distribution benchmarks show that Defake-o3 improves both detection accuracy and explanation quality, producing more localized, verifiable, and persuasive evidence.

AIGI检测可解释性证据验证多模态模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。