arXiv:2508.20182cs.CV2025-08

用Stable Diffusion模型直接定位图像伪造,无需大量标注数据。

SDiFL: Stable Diffusion-Driven Framework for Image Forgery Localization

  • 将伪造残留信号作为新模态注入SD3的隐空间进行训练。
  • 在多个基准数据集上性能提升最高达12%。
  • 无需训练数据中的伪造样本也能准确识别真实伪造图像。

受新一代多模态大模型(如Stable Diffusion, SD)驱动,图像篡改技术迅速发展,对图像取证构成严峻挑战。现有伪造定位方法严重依赖人工标注且成本高昂,难以跟上新技术步伐。为此,我们首次将SD的图像生成与强大感知能力整合进取证框架,实现更高效、精准的伪造定位。理论上证明,SD的多模态架构可基于伪造信息进行条件生成,从而内生输出定位结果。在此基础上,我们利用Stable Diffusion V3(SD3)的多模态能力,在隐空间中将通过特定高通滤波器提取的图像伪造残留(高频率信号)作为显式模态融合,增强定位性能。该方法完整保留SD3提取的隐特征,从而维持输入图像的丰富语义信息。实验表明,本框架在主流基准数据集上相较当前最优模型性能提升最高达12%。令人鼓舞的是,即使训练时未接触过真实文档伪造或自然场景篡改图像,模型仍表现出优异的泛化能力。

原文摘要 · Abstract (English)

Driven by the new generation of multi-modal large models, such as Stable Diffusion (SD), image manipulation technologies have advanced rapidly, posing significant challenges to image forensics. However, existing image forgery localization methods, which heavily rely on labor-intensive and costly annotated data, are struggling to keep pace with these emerging image manipulation technologies. To address these challenges, we are the first to integrate both image generation and powerful perceptual capabilities of SD into an image forensic framework, enabling more efficient and accurate forgery localization. First, we theoretically show that the multi-modal architecture of SD can be conditioned on forgery-related information, enabling the model to inherently output forgery localization results. Then, building on this foundation, we specifically leverage the multimodal framework of Stable DiffusionV3 (SD3) to enhance forgery localization performance.We leverage the multi-modal processing capabilities of SD3 in the latent space by treating image forgery residuals -- high-frequency signals extracted using specific highpass filters -- as an explicit modality. This modality is fused into the latent space during training to enhance forgery localization performance. Notably, our method fully preserves the latent features extracted by SD3, thereby retaining the rich semantic information of the input image. Experimental results show that our framework achieves up to 12% improvements in performance on widely used benchmarking datasets compared to current state-of-the-art image forgery localization models. Encouragingly, the model demonstrates strong performance on forensic tasks involving real-world document forgery images and natural scene forging images, even when such data were entirely unseen during training.

图像取证伪造定位Stable Diffusion隐空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。