用多模态大模型实现可解释的AI生成图像检测
Interpretable and Reliable Detection of AI-Generated Images via Grounded Reasoning in MLLMs
- 通过多阶段微调让大模型学会定位图像伪造痕迹
- 在检测准确率和解释一致性上显著优于现有方法
- 适合需要可信图像审核的媒体与安全领域
图像生成技术的快速发展催生了对可解释且可靠的检测方法的需求。现有方法虽精度高,但多为黑箱,缺乏人类可理解的解释。多模态大语言模型(MLLMs)虽非专为伪造检测设计,但具备强大的分析与推理能力。经适当微调后,能有效识别AI生成图像并提供有意义的解释。然而,现有MLLMs仍存在幻觉问题,难以将视觉解读与实际图像内容及人类认知对齐。为此,我们构建了一个包含边界框和描述性标题的AI生成图像数据集,突出合成伪影,为人类对齐的视觉-文本接地推理奠定基础。随后,采用多阶段优化策略微调MLLMs,逐步平衡检测准确性、视觉定位与文本解释连贯性的目标。结果表明,该模型在检测AI生成图像及定位视觉缺陷方面均表现优异,显著超越基线方法。
原文摘要 · Abstract (English)
The rapid advancement of image generation technologies intensifies the demand for interpretable and robust detection methods. Although existing approaches often attain high accuracy, they typically operate as black boxes without providing human-understandable justifications. Multi-modal Large Language Models (MLLMs), while not originally intended for forgery detection, exhibit strong analytical and reasoning capabilities. When properly fine-tuned, they can effectively identify AI-generated images and offer meaningful explanations. However, existing MLLMs still struggle with hallucination and often fail to align their visual interpretations with actual image content and human reasoning. To bridge this gap, we construct a dataset of AI-generated images annotated with bounding boxes and descriptive captions that highlight synthesis artifacts, establishing a foundation for human-aligned visual-textual grounded reasoning. We then finetune MLLMs through a multi-stage optimization strategy that progressively balances the objectives of accurate detection, visual localization, and coherent textual explanation. The resulting model achieves superior performance in both detecting AI-generated images and localizing visual flaws, significantly outperforming baseline methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。