arXiv:2412.17632cs.AIcs.CV2024-12中稿 · ACM MM 2025被引 2

构建数据集与评估框架,量化分析AI生成图像与真实图像的差距。

D-Judge: How Far Are We? Assessing the Discrepancies Between AI-synthesized and Natural Images through Multimodal Guidance

  • 构建包含44万张生成图的大规模多模态数据集D-ANI
  • 在5个维度上发现AI图像与真实图像存在显著差异
  • 为评估生成图像质量提供可量化的基准,适合模型开发者参考

在快速发展的人工智能生成内容(AIGC)领域,区分AI生成图像与自然图像仍是核心挑战。尽管先进生成模型能产出视觉逼真的图像,但与真实图像仍存在显著差异。为此,我们构建了大规模多模态数据集D-ANI,包含5,000张自然图像和超过440,000张由九种代表性模型生成的AIGI样本,覆盖文本到图像(T2I)、图像到图像(I2I)及文本与图像到图像(TI2I)等多种提示方式。我们提出AI-自然图像差异评估基准D-Judge,系统回答‘AI生成图像离真实图像有多远’这一关键问题。通过五个维度——直观视觉质量、语义对齐度、审美吸引力、下游任务适用性及人工协同验证——对D-ANI进行细粒度评估。大量实验揭示各维度均存在显著差距,凸显将量化指标与人类判断对齐的重要性,以全面理解AI生成图像质量。代码:https://github.com/ryliu68/DJudge;数据:https://huggingface.co/datasets/Renyang/DANI。

原文摘要 · Abstract (English)

In the rapidly evolving field of Artificial Intelligence Generated Content (AIGC), a central challenge is distinguishing AI-synthesized images from natural ones. Despite the impressive capabilities of advanced generative models in producing visually compelling images, significant discrepancies remain when compared to natural images. To systematically investigate and quantify these differences, we construct a large-scale multimodal dataset, D-ANI, comprising 5,000 natural images and over 440,000 AIGI samples generated by nine representative models using both unimodal and multimodal prompts, including Text-to-Image (T2I), Image-to-Image (I2I), and Text-and-Image-to-Image (TI2I). We then introduce an AI-Natural Image Discrepancy assessment benchmark (D-Judge) to address the critical question: how far are AI-generated images (AIGIs) from truly realistic images? Our fine-grained evaluation framework assesses the D-ANI dataset across five dimensions: naive visual quality, semantic alignment, aesthetic appeal, downstream task applicability, and coordinated human validation. Extensive experiments reveal substantial discrepancies across these dimensions, highlighting the importance of aligning quantitative metrics with human judgment to achieve a comprehensive understanding of AI-generated image quality. Code: https://github.com/ryliu68/DJudge ; Data: https://huggingface.co/datasets/Renyang/DANI.

图像生成质量评估多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。