arXiv:2509.10260cs.CV2025-09被引 4

构建首个大规模细粒度图像生成瑕疵评估数据集与基准。

MagicMirror: A Large-Scale Dataset and Benchmark for Fine-Grained Artifacts Assessment in Text-to-Image Generation

  • 提出细粒度瑕疵分类体系,人工标注34万张生成图像。
  • 训练VLM模型精准识别并分类图像中的结构与生理缺陷。
  • 发现顶级模型仍普遍存在显著瑕疵,适合评测生成质量研究者使用。

文本到图像(T2I)生成在指令遵循和美学表现上取得显著进展,但物理瑕疵(如解剖或结构错误)仍广泛存在,严重降低感知质量并限制实际应用。由于这些瑕疵类型多样且复杂,现有评估基准缺乏系统性与细粒度。为此,我们提出MagicMirror框架:首先建立生成图像瑕疵的详细分类体系;在此指导下,手动标注了首个大规模人类标注数据集MagicData340K,包含34万张带细粒度标签的生成图像;基于该数据集,训练了MagicAssessor——一个能提供详细评估与标签的视觉-语言模型(VLM);为应对类别不平衡与奖励劫持问题,设计新型数据采样策略与多级奖励系统,用于组相对策略优化(GRPO);最终利用MagicAssessor构建自动化基准MagicBench,用于评估当前T2I模型的图像瑕疵。对主流模型的评估显示,即使顶级模型如GPT-image-1也持续存在显著瑕疵,凸显减少生成瑕疵是未来T2I发展的关键前沿。

原文摘要 · Abstract (English)

Text-to-image (T2I) generation has achieved remarkable progress in instruction following and aesthetics. However, a persistent challenge is the prevalence of physical artifacts, such as anatomical and structural flaws, which severely degrade perceptual quality and limit application. Given the diversity and complexity of these artifacts, a systematic and fine-grained evaluation framework is required, which is lacking in current benchmarks. To fill this gap, we introduce MagicMirror, a comprehensive framework for artifacts assessment. We first establish a detailed taxonomy of generated image artifacts. Guided by this taxonomy, we manually annotate MagicData340K, the first human-annotated large-scale dataset of 340K generated images with fine-grained artifact labels. Building on this dataset, we train MagicAssessor, a Vision-Language Model (VLM) that provides detailed assessments and corresponding labels. To overcome challenges like class imbalance and reward hacking, we design a novel data sampling strategy and a multi-level reward system for Group Relative Policy Optimization (GRPO). Finally, we leverage MagicAssessor to construct MagicBench, an automated benchmark for evaluating the image artifacts of current T2I models. Our evaluation with MagicBench reveals that despite their widespread adoption, even top-tier models like GPT-image-1 are consistently plagued by significant artifacts, highlighting artifact reduction as a critical frontier for future T2I development. Project page: https://wj-inf.github.io/MagicMirror-page/.

图像生成瑕疵评估大模型评测VLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。