用多模态大模型实现可解释的图像伪造检测与定位
FakeShield: Explainable Image Forgery Detection and Localization via Multi-modal Large Language Models
- 结合图文分析,通过像素级和图像级线索判断真伪
- 在多种篡改类型上表现优异,定位精度显著提升
- 适合需要透明决策过程的安全审核场景
生成式AI的快速发展带来内容创作便利的同时,也使图像篡改更易实现且难以识别。现有图像伪造检测与定位(IFDL)方法虽有效,但存在黑箱决策和跨篡改类型泛化能力差的问题。为此,我们提出可解释的IFDL任务,并设计FakeShield多模态框架,具备评估图像真实性、生成篡改区域掩码及提供基于像素级和图像级线索的判据能力。我们利用GPT-4o增强现有IFDL数据集,构建多模态篡改描述数据集MMTD-Set以训练模型。引入领域标签引导的可解释伪造检测模块(DTE-FDM)和多模态伪造定位模块(MFLM),实现基于详细文本描述的伪造定位。大量实验表明,FakeShield能有效检测并定位多种篡改技术,相比以往方法更具可解释性和优越性。
原文摘要 · Abstract (English)
The rapid development of generative AI is a double-edged sword, which not only facilitates content creation but also makes image manipulation easier and more difficult to detect. Although current image forgery detection and localization (IFDL) methods are generally effective, they tend to face two challenges: \textbf{1)} black-box nature with unknown detection principle, \textbf{2)} limited generalization across diverse tampering methods (e.g., Photoshop, DeepFake, AIGC-Editing). To address these issues, we propose the explainable IFDL task and design FakeShield, a multi-modal framework capable of evaluating image authenticity, generating tampered region masks, and providing a judgment basis based on pixel-level and image-level tampering clues. Additionally, we leverage GPT-4o to enhance existing IFDL datasets, creating the Multi-Modal Tamper Description dataSet (MMTD-Set) for training FakeShield's tampering analysis capabilities. Meanwhile, we incorporate a Domain Tag-guided Explainable Forgery Detection Module (DTE-FDM) and a Multi-modal Forgery Localization Module (MFLM) to address various types of tamper detection interpretation and achieve forgery localization guided by detailed textual descriptions. Extensive experiments demonstrate that FakeShield effectively detects and localizes various tampering techniques, offering an explainable and superior solution compared to previous IFDL methods. The code is available at https://github.com/zhipeixu/FakeShield.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。