arXiv:2601.00553cs.CVcs.AI2026-01被引 5

构建9.6万张图像数据集,用于区分真实与AI生成图片。

A Comprehensive Dataset for Human vs. AI Generated Image Detection

  • 基于MS COCO构建数据集,整合5种主流生成模型。
  • 包含96000张图像,支持真伪分类与生成模型溯源任务。
  • 适合媒体安全、内容审核研究者使用。

Stable Diffusion、DALL-E 和 MidJourney 等多模态生成模型彻底改变了合成图像的生成方式。这些工具推动创新的同时,也助长了误导性内容、虚假信息和伪造媒体的传播。随着生成图像越来越接近真实照片,检测其来源已成为当务之急。为此,我们发布了 MS COCOAI,一个全新的用于 AI 生成图像检测的数据集,包含 96000 个真实与合成图像样本,基于 MS COCO 构建。合成图像由五种生成器创建:Stable Diffusion 3、Stable Diffusion 2.1、SDXL、DALL-E 3 和 MidJourney v6。基于该数据集,我们提出两项任务:(1)将图像分类为真实或生成;(2)识别给定合成图像所使用的生成模型。数据集可在 https://huggingface.co/datasets/Rajarshi-Roy-research/Defactify_Image_Dataset 获取。

原文摘要 · Abstract (English)

Multimodal generative AI systems like Stable Diffusion, DALL-E, and MidJourney have fundamentally changed how synthetic images are created. These tools drive innovation but also enable the spread of misleading content, false information, and manipulated media. As generated images become harder to distinguish from photographs, detecting them has become an urgent priority. To combat this challenge, we release MS COCOAI, a novel dataset for AI generated image detection consisting of 96000 real and synthetic datapoints, built using the MS COCO dataset. To generate synthetic images, we use five generators: Stable Diffusion 3, Stable Diffusion 2.1, SDXL, DALL-E 3, and MidJourney v6. Based on the dataset, we propose two tasks: (1) classifying images as real or generated, and (2) identifying which model produced a given synthetic image. The dataset is available at https://huggingface.co/datasets/Rajarshi-Roy-research/Defactify_Image_Dataset.

图像检测数据集AI生成内容安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。