arXiv:2603.12930cs.CVcs.LG2026-03被引 1

用视觉语言模型提升图像伪造检测与定位,效果更准且结果更可解释。

Rethinking VLMs for Image Forgery Detection and Localization

  • 利用伪造区域掩码作为额外先验,引导模型学习真实特征
  • 在9个基准上达到新最佳性能,跨数据集泛化能力更强
  • 适合关注伪造内容检测与模型可解释性的研究人员

随着人工智能生成内容(AIGC)的快速发展,图像篡改日益普及,给图像伪造检测与定位(IFDL)带来严峻挑战。本文研究如何充分挖掘视觉语言模型(VLMs)在该任务中的潜力。我们发现,传统VLMs依赖的语义合理性先验对检测与定位帮助有限,甚至因固有偏见产生负面效果。相反,显式编码伪造概念的位置掩码可作为额外先验,辅助VLM训练优化,提升结果可解释性。基于此,我们提出新框架IFDL-VLM。在9个主流基准上进行实验,涵盖域内与跨数据集泛化场景,结果表明该方法在检测、定位与可解释性方面均实现新的最先进性能。代码已开源。

原文摘要 · Abstract (English)

With the rapid rise of Artificial Intelligence Generated Content (AIGC), image manipulation has become increasingly accessible, posing significant challenges for image forgery detection and localization (IFDL). In this paper, we study how to fully leverage vision-language models (VLMs) to assist the IFDL task. In particular, we observe that priors from VLMs hardly benefit the detection and localization performance and even have negative effects due to their inherent biases toward semantic plausibility rather than authenticity. Additionally, the location masks explicitly encode the forgery concepts, which can serve as extra priors for VLMs to ease their training optimization, thus enhancing the interpretability of detection and localization results. Building on these findings, we propose a new IFDL pipeline named IFDL-VLM. To demonstrate the effectiveness of our method, we conduct experiments on 9 popular benchmarks and assess the model performance under both in-domain and cross-dataset generalization settings. The experimental results show that we consistently achieve new state-of-the-art performance in detection, localization, and interpretability.Code is available at: https://github.com/sha0fengGuo/IFDL-VLM.

图像伪造视觉语言模型检测定位可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。