构建4000张生成图像缺陷数据集,支持缺陷检测、定位与解释。
AGIDefect-4K: A Richly Annotated Dataset for AI-Generated Image Defect Detection, Localization and Explanation

- 构建15种生成模型的4000张图像缺陷数据集,含多层级标注。
- 提出AGIDA框架,实现缺陷检测、定位、解释与质量评分联合预测。
- 适合研究生成图像可靠性、可解释性与评测基准的学者使用。
生成式AI已能产出高度逼真的图像,但现有模型仍存在细微却关键的缺陷,影响其可靠性。尽管已有评估基准取得进展,对生成图像缺陷的全面诊断仍缺乏系统研究。为此,我们推出AGIDefect-4K数据集,包含来自15个前沿生成模型(涵盖开源与闭源)的4,000张图像,提供多层次标注:(1) 缺陷是否存在,(2) 像素级分割掩码定位缺陷区域,(3) 文本化解释描述缺陷类型及其感知影响,并为每张图附带整体质量评分。基于此,我们提出AGIDA(AGI缺陷助手)框架,利用多模态大语言模型(MLLMs)实现缺陷检测、定位、解释与质量预测的联合建模。在AGIDefect-4K上的全面评测表明,当前对生成图像缺陷的理解仍具挑战性,凸显该数据集的价值。数据集已公开于 https://github.com/sxfly99/AGIDefect-4K。
原文摘要 · Abstract (English)
Generative AI can now produce highly realistic images, yet current models still exhibit subtle but critical defects that undermine their reliability. While existing AI-generated image (AGI) evaluation benchmarks have made notable progress, comprehensive AGI defect diagnosis remains underexplored. To bridge this gap, we introduce AGIDefect-4K, a richly annotated dataset of 4,000 images from 15 state-of-the-art generative models spanning both open-source and closed-source systems. AGIDefect-4K features hierarchical defect annotations: (1) detection labels identifying whether defects exist, (2) pixel-level segmentation masks localizing defective regions, and (3) detailed textual explanations characterizing defect types and their perceptual impact. Each image is further annotated with an overall quality score. Building on this, we present AGIDA (AGI Defect Assistant), a baseline framework leveraging Multimodal Large Language Models (MLLMs) for joint defect detection, localization, explanation, and quality prediction. Comprehensive benchmarking on AGIDefect-4K reveals that AGI defect understanding remains challenging, underscoring the value of this dataset. The dataset is publicly available at https://github.com/sxfly99/AGIDefect-4K.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。