用多模态大模型定位扩散模型编辑的伪造区域,提升图像鉴伪能力。
EditScout: Locating Forged Regions from Diffusion-based Edited Images with Multimodal LLM
- 融合多模态大模型,利用语义理解增强伪造区域定位能力。
- 在MagicBrush、AutoSplice和PerfBrush数据集上mIoU和F1-score均领先。
- 对新型编辑手法尤其有效,适合数字取证与媒体可信度研究者。
图像编辑技术可对图像进行变换、调整、删除等操作。近期研究显著提升了图像编辑工具的能力,能生成几乎无法与真实图像区分的逼真且语义合理的伪造区域,给数字取证和媒体可信性带来新挑战。当前图像取证技术虽擅长定位传统图像处理产生的伪造区域,但在应对基于扩散模型的编辑方法时表现不佳。为此,我们提出一种新框架,整合多模态大型语言模型(LLM)以增强推理能力,实现对扩散模型生成编辑图像中篡改区域的定位。通过利用LLM的上下文与语义优势,该框架在MagicBrush、AutoSplice和PerfBrush(新型扩散模型数据集)上取得良好效果,优于以往方法,在mIoU和F1-score指标上表现更优。尤其在PerfBrush数据集——一个包含此前未见编辑类型的自建测试集上,传统方法普遍表现差、得分极低,而本方法展现出显著性能优势。
原文摘要 · Abstract (English)
Image editing technologies are tools used to transform, adjust, remove, or otherwise alter images. Recent research has significantly improved the capabilities of image editing tools, enabling the creation of photorealistic and semantically informed forged regions that are nearly indistinguishable from authentic imagery, presenting new challenges in digital forensics and media credibility. While current image forensic techniques are adept at localizing forged regions produced by traditional image manipulation methods, current capabilities struggle to localize regions created by diffusion-based techniques. To bridge this gap, we present a novel framework that integrates a multimodal Large Language Model (LLM) for enhanced reasoning capabilities to localize tampered regions in images produced by diffusion model-based editing methods. By leveraging the contextual and semantic strengths of LLMs, our framework achieves promising results on MagicBrush, AutoSplice, and PerfBrush (novel diffusion-based dataset) datasets, outperforming previous approaches in mIoU and F1-score metrics. Notably, our method excels on the PerfBrush dataset, a self-constructed test set featuring previously unseen types of edits. Here, where traditional methods typically falter, achieving markedly low scores, our approach demonstrates promising performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。