用多模态大模型实现可解释的假图检测,让判断过程透明可信。
Towards Explainable Fake Image Detection with Multi-Modal Large Language Models
- 利用多模态大模型进行基于推理的假图检测
- 相比传统方法和人类,模型在跨域泛化上表现更优
- 设计六种提示框架,提升检测可解释性,适合安全审查场景
图像生成技术的发展带来严峻的公共安全挑战。我们认为,假图检测不应是‘黑箱’操作,理想的方案必须兼具强泛化能力与透明性。多模态大语言模型(MLLMs)为基于推理的AI生成图像检测提供了新机遇。本文评估了MLLMs相较于传统检测方法和人类评估者的性能,揭示其优势与局限。进一步地,我们设计了六种不同提示策略,并提出一个集成框架,构建出更鲁棒、可解释且以推理为核心的检测系统。代码已开源:https://github.com/Gennadiyev/mllm-defake。
原文摘要 · Abstract (English)
Progress in image generation raises significant public security concerns. We argue that fake image detection should not operate as a "black box". Instead, an ideal approach must ensure both strong generalization and transparency. Recent progress in Multi-modal Large Language Models (MLLMs) offers new opportunities for reasoning-based AI-generated image detection. In this work, we evaluate the capabilities of MLLMs in comparison to traditional detection methods and human evaluators, highlighting their strengths and limitations. Furthermore, we design six distinct prompts and propose a framework that integrates these prompts to develop a more robust, explainable, and reasoning-driven detection system. The code is available at https://github.com/Gennadiyev/mllm-defake.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。