arXiv:2505.15644cs.CVcs.AI2025-05被引 4

首次系统评估视觉语言模型对精细图像编辑的检测与定位能力。

Can VLMs Detect and Localize Fine-Grained AI-Edited Images?

  • 构建自动化数据生成流程,创建大规模编辑图像基准FragFake。
  • 微调后的VLM如Qwen2.5-VL在定位精度上显著优于预训练模型。
  • 揭示跨编辑器、跨数据集泛化能力的潜力与局限,适合内容真实性研究者。

细粒度检测与定位局部图像编辑对评估内容真实性至关重要,尤其在现代扩散模型和图像编辑工具能生成高度逼真的篡改内容背景下。然而,该任务面临三大挑战:(1) 多数AIGC检测器仅输出全局真/假标签,无法指示编辑位置;(2) 传统计算机视觉方法依赖昂贵的像素级标注;(3) 缺乏大规模、面向现代编辑场景的基准数据集。为此,我们开发了自动化数据生成流程,构建了FragFake——一个涵盖多个源数据集、多样编辑模型和常见编辑类型的大型基准。基于此,我们首次系统研究视觉语言模型(VLMs)在图像编辑分类与区域定位中的表现。实验表明,预训练VLM(如GPT4o)表现不佳,而微调模型如Qwen2.5-VL在各类设置下均实现高准确率与更高物体精确率。进一步探索基于GRPO的强化学习可微调视觉推理(RLVR)训练,带来适度性能提升并增强模型输出可解释性。消融与迁移分析揭示数据平衡、训练规模、LoRA秩及训练域对性能的影响,凸显跨编辑器与跨数据集泛化的潜力与局限。本工作有望为多模态内容真实性研究奠定坚实基础。

原文摘要 · Abstract (English)

Fine-grained detection and localization of localized image edits is crucial for assessing content authenticity, especially as modern diffusion models and image editors can produce highly realistic manipulations. However, this problem faces three key challenges: (1) most AIGC detectors produce only a global real-or-fake label without indicating where edits occur; (2) traditional computer vision methods for edit localization typically rely on costly pixel-level annotations; and (3) there is no large-scale, modern benchmark specifically targeting edited-image detection. To address these gaps, we develop an automated data-generation pipeline and construct FragFake, a large-scale benchmark of AI-edited images spanning multiple source datasets, diverse editing models, and several common edit types. Building on FragFake, we are the first to systematically study vision language models (VLMs) for edited-image classification and edited-region localization. Our experiments show that pretrained VLMs, including GPT4o, perform poorly on this task, whereas fine-tuned models such as Qwen2.5-VL achieve high accuracy and substantially higher object precision across all settings. We further explore GRPO-based RLVR training, which yields modest metric gains while improving the interpretability of model outputs. Ablation and transfer analyses reveal how data balancing, training size, LoRA rank, and training domain affect performance, and highlight both the potential and the limitations of cross-editor and cross-dataset generalization. We anticipate that this work will establish a solid foundation to facilitate and inspire subsequent research endeavors in the domain of multimodal content authenticity.

图像伪造检测视觉语言模型内容真实性定位任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。