用大模型检测图像篡改,能定位问题区域还避免胡编乱造。
ForgerySleuth: Empowering Multimodal Large Language Models for Image Manipulation Detection
- 利用大模型融合多线索,生成篡改区域分割图。
- 在新数据集上表现超越现有方法,泛化与可解释性更强。
- 适合需要精准、可信图像验证的场景或研究者。
多模态大语言模型在各类多模态任务中展现出巨大潜力,但在图像篡改检测(IMD)任务中的应用尚未被充分探索。直接应用于该任务时,M-LLMs常产生存在幻觉和过度推理的分析文本。为此,我们提出 ForgerySleuth,通过 M-LLMs 实现全面线索融合,并生成指示具体篡改区域的分割输出。此外,我们基于 Chain-of-Clues 提示构建了 ForgeryAnalysis 数据集,包含分析与推理文本,以提升图像篡改检测任务质量。同时引入数据引擎,为预训练阶段构建更大规模数据集。大量实验表明,ForgeryAnalysis 有效,且 ForgerySleuth 在泛化能力、鲁棒性和可解释性方面显著优于现有方法。
原文摘要 · Abstract (English)
Multimodal large language models have unlocked new possibilities for various multimodal tasks. However, their potential in image manipulation detection remains unexplored. When directly applied to the IMD task, M-LLMs often produce reasoning texts that suffer from hallucinations and overthinking. To address this, we propose ForgerySleuth, which leverages M-LLMs to perform comprehensive clue fusion and generate segmentation outputs indicating specific regions that are tampered with. Moreover, we construct the ForgeryAnalysis dataset through the Chain-of-Clues prompt, which includes analysis and reasoning text to upgrade the image manipulation detection task. A data engine is also introduced to build a larger-scale dataset for the pre-training phase. Our extensive experiments demonstrate the effectiveness of ForgeryAnalysis and show that ForgerySleuth significantly outperforms existing methods in generalization, robustness, and explainability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。