arXiv:2411.04125cs.CV2024-11CVPR被引 57

用4803个生成模型训练假图检测器,提升泛化能力。

Community Forensics: Using Thousands of Generators to Train Fake Image Detectors

  • 从4803个模型中采样270万张图像构建新数据集
  • 模型数量越多、多样性越高,检测效果越好
  • 适合关注生成图像检测泛化能力的研究者

检测由未知生成模型创建的AI伪造图像是一大挑战。我们指出,训练数据多样性不足是主要障碍,并提出一个显著更大更丰富的数据集。为构建该数据集,我们系统性地下载数千个文本到图像的潜在扩散模型并从中采样图像,同时收集数十个主流开源与商业模型的图像。最终数据集包含270万张图像,源自4803个不同模型,涵盖广泛的场景内容、生成架构和图像处理设置。利用该数据集,我们研究了假图检测器的泛化能力。实验表明,即使模型架构相似,训练集中模型数量增加也会提升检测性能;模型多样性提高时,检测性能也随之提升,且我们的检测器在泛化能力上优于基于其他数据集训练的模型。数据集可访问:https://jespark.net/projects/2024/community_forensics

原文摘要 · Abstract (English)

One of the key challenges of detecting AI-generated images is spotting images that have been created by previously unseen generative models. We argue that the limited diversity of the training data is a major obstacle to addressing this problem, and we propose a new dataset that is significantly larger and more diverse than prior work. As part of creating this dataset, we systematically download thousands of text-to-image latent diffusion models and sample images from them. We also collect images from dozens of popular open source and commercial models. The resulting dataset contains 2.7M images that have been sampled from 4803 different models. These images collectively capture a wide range of scene content, generator architectures, and image processing settings. Using this dataset, we study the generalization abilities of fake image detectors. Our experiments suggest that detection performance improves as the number of models in the training set increases, even when these models have similar architectures. We also find that detection performance improves as the diversity of the models increases, and that our trained detectors generalize better than those trained on other datasets. The dataset can be found in https://jespark.net/projects/2024/community_forensics

图像伪造检测生成模型数据集构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。