arXiv:2503.15264cs.CV2025-03ICCV被引 71

用多模态大模型实现伪造图像的精准定位与可解释检测。

LEGION: Learning to Ground and Explain for Synthetic Image Detection

  • 基于多模态大模型,融合检测、分割与文本解释能力。
  • 在合成图像数据集上,检测精度超越传统专家3.31%(mIoU)。
  • 可作为生成控制模块,提升图像真实感,契合人类偏好。

生成技术的快速发展带来便利的同时也引发社会担忧。现有伪造图像检测方法普遍缺乏像素级可解释性,且过度关注图像篡改检测,而数据集常因生成器过时和标注粗糙而受限。本文提出高质量、多样化的合成图像数据集 SynthScars,包含12,236张完全合成图像,涵盖4类图像内容、3类伪造痕迹,并提供像素级分割、详细文本解释及痕迹类别标签等细粒度标注。同时,我们提出 LEGION(LEarning to Ground and explain for Synthetic Image detectiON),一种基于多模态大语言模型的伪造分析框架,集成伪造检测、分割与解释功能。进一步地,我们将 LEGION 作为控制器,嵌入图像修复流程中,引导生成更真实、更高质量的图像。大量实验表明,LEGION 在多个基准测试中优于现有方法,尤其在 SynthScars 上比第二佳的传统专家高出3.31%(mIoU)和7.75%(F1)。经其指导生成的图像与人类偏好更具一致性。代码、模型与数据集将公开发布。

原文摘要 · Abstract (English)

The rapid advancements in generative technology have emerged as a double-edged sword. While offering powerful tools that enhance convenience, they also pose significant social concerns. As defenders, current synthetic image detection methods often lack artifact-level textual interpretability and are overly focused on image manipulation detection, and current datasets usually suffer from outdated generators and a lack of fine-grained annotations. In this paper, we introduce SynthScars, a high-quality and diverse dataset consisting of 12,236 fully synthetic images with human-expert annotations. It features 4 distinct image content types, 3 categories of artifacts, and fine-grained annotations covering pixel-level segmentation, detailed textual explanations, and artifact category labels. Furthermore, we propose LEGION (LEarning to Ground and explain for Synthetic Image detectiON), a multimodal large language model (MLLM)-based image forgery analysis framework that integrates artifact detection, segmentation, and explanation. Building upon this capability, we further explore LEGION as a controller, integrating it into image refinement pipelines to guide the generation of higher-quality and more realistic images. Extensive experiments show that LEGION outperforms existing methods across multiple benchmarks, particularly surpassing the second-best traditional expert on SynthScars by 3.31% in mIoU and 7.75% in F1 score. Moreover, the refined images generated under its guidance exhibit stronger alignment with human preferences. The code, model, and dataset will be released.

伪造检测可解释性多模态图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。