arXiv:2511.13442cs.CVcs.AI2025-11

无需训练,用通用大模型高效识别图像伪造并给出详细解释

Unlocking the Forgery Detection Potential of Vanilla MLLMs: A Novel Training-Free Pipeline

  • 不需额外训练,直接利用大模型原生能力分析图像伪造
  • 在多种篡改类型上定位精度超越现有方法,解释更全面
  • 适合需要快速部署、解释性强的图像真实性检测场景

随着人工智能生成内容(AIGC)技术的发展,包括多模态大语言模型(MLLM)和扩散模型在内的工具使图像生成与编辑变得极为便捷。现有图像伪造检测与定位(IFDL)方法普遍泛化能力弱,且解释性不足。尽管当前研究尝试将大模型引入该任务,但依赖大规模训练,计算成本高,且未能充分挖掘原始大模型的潜力。为此,我们提出Foresee——一种无需训练的基于大模型的图像伪造分析新流程。该方法通过类型先验驱动策略与灵活特征检测模块(FFD),专门应对复制-粘贴篡改,有效释放了原生大模型在取证领域的潜力。实验表明,Foresee在多种篡改类型(包括复制-粘贴、拼接、删除、局部增强、深度伪造及AIGC编辑)上均实现更高定位精度,并提供更丰富的文本解释,同时具备更强泛化能力,显著优于现有方法。代码将在最终版本发布。

原文摘要 · Abstract (English)

With the rapid advancement of artificial intelligence-generated content (AIGC) technologies, including multimodal large language models (MLLMs) and diffusion models, image generation and manipulation have become remarkably effortless. Existing image forgery detection and localization (IFDL) methods often struggle to generalize across diverse datasets and offer limited interpretability. Nowadays, MLLMs demonstrate strong generalization potential across diverse vision-language tasks, and some studies introduce this capability to IFDL via large-scale training. However, such approaches cost considerable computational resources, while failing to reveal the inherent generalization potential of vanilla MLLMs to address this problem. Inspired by this observation, we propose Foresee, a training-free MLLM-based pipeline tailored for image forgery analysis. It eliminates the need for additional training and enables a lightweight inference process, while surpassing existing MLLM-based methods in both tamper localization accuracy and the richness of textual explanations. Foresee employs a type-prior-driven strategy and utilizes a Flexible Feature Detector (FFD) module to specifically handle copy-move manipulations, thereby effectively unleashing the potential of vanilla MLLMs in the forensic domain. Extensive experiments demonstrate that our approach simultaneously achieves superior localization accuracy and provides more comprehensive textual explanations. Moreover, Foresee exhibits stronger generalization capability, outperforming existing IFDL methods across various tampering types, including copy-move, splicing, removal, local enhancement, deepfake, and AIGC-based editing. The code will be released in the final version.

图像伪造检测大模型应用无训练方法可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。