arXiv:2511.19111cs.CV2025-11被引 1

构建3万张局部扩散编辑图像数据集,推动AI生成内容的精准定位检测。

DiffSeg30k: A Multi-Turn Diffusion Editing Benchmark for Localized AIGC Detection

  • 基于8种主流扩散模型,对真实图像进行多轮局部编辑
  • 每张图含像素级标注,支持细粒度检测与编辑模型识别
  • 适合研究生成内容溯源、图像篡改检测的学者和工程师

基于扩散模型的局部编辑使图像修改更逼真,增加了AI生成内容检测的难度。现有检测基准多聚焦整图分类,忽视了编辑位置的定位。本文提出DiffSeg30k,一个公开可获取的3万张扩散编辑图像数据集,配备像素级标注,支持细粒度检测。该数据集包含:1)来自COCO的真实场景图像或提示;2)使用8种前沿扩散模型进行局部编辑;3)每张图像最多经历三次连续编辑,模拟真实操作流程;4)通过视觉-语言模型(VLM)自动识别有意义区域,并生成上下文感知的编辑提示,覆盖增删与属性修改。该数据集将AIGC检测从二分类转向语义分割,实现编辑位置定位与编辑模型识别的同步。我们评估三种基础分割方法,揭示其在应对图像失真时的显著挑战。实验还发现,尽管训练目标是像素级定位,这些分割模型在整图分类上表现优异,超越传统伪造检测器,且具备良好的跨生成器泛化能力。我们认为DiffSeg30k将推动细粒度生成内容定位的研究,展现分割方法的潜力与局限。数据集已发布于https://huggingface.co/datasets/Chaos2629/Diffseg30k。

原文摘要 · Abstract (English)

Diffusion-based editing enables realistic modification of local image regions, making AI-generated content harder to detect. Existing AIGC detection benchmarks focus on classifying entire images, overlooking the localization of diffusion-based edits. We introduce DiffSeg30k, a publicly available dataset of 30k diffusion-edited images with pixel-level annotations, designed to support fine-grained detection. DiffSeg30k features: 1) In-the-wild images--we collect images or image prompts from COCO to reflect real-world content diversity; 2) Diverse diffusion models--local edits using eight SOTA diffusion models; 3) Multi-turn editing--each image undergoes up to three sequential edits to mimic real-world sequential editing; and 4) Realistic editing scenarios--a vision-language model (VLM)-based pipeline automatically identifies meaningful regions and generates context-aware prompts covering additions, removals, and attribute changes. DiffSeg30k shifts AIGC detection from binary classification to semantic segmentation, enabling simultaneous localization of edits and identification of the editing models. We benchmark three baseline segmentation approaches, revealing significant challenges in semantic segmentation tasks, particularly concerning robustness to image distortions. Experiments also reveal that segmentation models, despite being trained for pixel-level localization, emerge as highly reliable whole-image classifiers of diffusion edits, outperforming established forgery classifiers while showing great potential in cross-generator generalization. We believe DiffSeg30k will advance research in fine-grained localization of AI-generated content by demonstrating the promise and limitations of segmentation-based methods. DiffSeg30k is released at: https://huggingface.co/datasets/Chaos2629/Diffseg30k

AIGC检测扩散模型图像编辑细粒度定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。