arXiv:2505.11257cs.CVcs.LG2025-05被引 4

构建了涵盖25个扩散模型的高质量合成图像数据集,助力假图检测研究。

DRAGON: A Large-Scale Dataset of Realistic Images Generated by Diffusion Models

  • 用大语言模型扩增提示词,提升生成图像多样性与真实感
  • 包含25个模型、多尺寸数据集,覆盖新旧主流架构
  • 专为伪造图像检测与溯源技术评估设计,提供基准测试集

扩散模型在图像生成中应用便捷,但也催生了大量虚假内容。为应对这一挑战,亟需高效可靠的检测工具。然而现有方法依赖大量训练样本,且多数数据集仅覆盖有限模型,迅速过时。本文提出DRAGON,一个涵盖25个扩散模型的综合性数据集,覆盖从经典到前沿的多种架构,包含丰富多样的图像内容。为提升图像真实感,我们设计了一种简单有效的流水线:利用大语言模型扩展输入提示,显著提升生成图像质量(在标准指标上表现更优)。数据集提供从极小到超大规模的多种版本,适配不同研究需求。同时配套专用测试集,可作为新型检测方法的基准评估平台。该数据集旨在支持合成内容的取证研究,推动检测与溯源技术发展。

原文摘要 · Abstract (English)

The remarkable ease of use of diffusion models for image generation has led to a proliferation of synthetic content online. While these models are often employed for legitimate purposes, they are also used to generate fake images that support misinformation and hate speech. Consequently, it is crucial to develop robust tools capable of detecting whether an image has been generated by such models. Many current detection methods, however, require large volumes of sample images for training. Unfortunately, due to the rapid evolution of the field, existing datasets often cover only a limited range of models and quickly become outdated. In this work, we introduce DRAGON, a comprehensive dataset comprising images from 25 diffusion models, spanning both recent advancements and older, well-established architectures. The dataset contains a broad variety of images representing diverse subjects. To enhance image realism, we propose a simple yet effective pipeline that leverages a large language model to expand input prompts, thereby generating more diverse and higher-quality outputs, as evidenced by improvements in standard quality metrics. The dataset is provided in multiple sizes (ranging from extra-small to extra-large) to accomodate different research scenarios. DRAGON is designed to support the forensic community in developing and evaluating detection and attribution techniques for synthetic content. Additionally, the dataset is accompanied by a dedicated test set, intended to serve as a benchmark for assessing the performance of newly developed methods.

扩散模型图像生成数据集假图检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。