arXiv:2505.19415cs.CV2025-05NeurIPS被引 17

构建首个综合多模态图像生成评估基准,统一测试指令遵循与图像一致性。

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models

  • 提出多层级评估框架,涵盖视觉瑕疵、语义对齐与美学偏好
  • 覆盖380个主体、4850条带标注提示,支持跨模态条件生成评测
  • 基于3.2万次人工评分验证,揭示模型架构与数据设计影响

当前多模态图像生成模型如GPT-4o、Gemini 2.0 Flash和Gemini 2.5 Pro在复杂指令理解、图像编辑和概念一致性方面表现优异,但现有评估工具割裂:文本到图像(T2I)基准缺乏多模态条件,定制化基准忽略组合语义与常识。本文提出MMIG-Bench,一个综合性多模态图像生成评估基准,通过4,850条丰富标注的文本提示与1,750张多视角参考图像(覆盖380个主体,包括人物、动物、物体及艺术风格)实现任务统一。该基准配备三级评估体系:(1) 低层指标评估视觉伪影与对象身份保持;(2) 新提出的方面匹配分数(AMS),基于VQA的中层指标,实现细粒度提示-图像对齐,与人类判断高度相关;(3) 高层指标评估美学与人类偏好。利用该基准,我们评测了17个前沿模型,包括Gemini 2.5 Pro、FLUX、DreamBooth和IP-Adapter,通过32,000次人工评分验证指标有效性,深入揭示模型架构与数据设计的影响。

原文摘要 · Abstract (English)

Recent multimodal image generators such as GPT-4o, Gemini 2.0 Flash, and Gemini 2.5 Pro excel at following complex instructions, editing images and maintaining concept consistency. However, they are still evaluated by disjoint toolkits: text-to-image (T2I) benchmarks that lacks multi-modal conditioning, and customized image generation benchmarks that overlook compositional semantics and common knowledge. We propose MMIG-Bench, a comprehensive Multi-Modal Image Generation Benchmark that unifies these tasks by pairing 4,850 richly annotated text prompts with 1,750 multi-view reference images across 380 subjects, spanning humans, animals, objects, and artistic styles. MMIG-Bench is equipped with a three-level evaluation framework: (1) low-level metrics for visual artifacts and identity preservation of objects; (2) novel Aspect Matching Score (AMS): a VQA-based mid-level metric that delivers fine-grained prompt-image alignment and shows strong correlation with human judgments; and (3) high-level metrics for aesthetics and human preference. Using MMIG-Bench, we benchmark 17 state-of-the-art models, including Gemini 2.5 Pro, FLUX, DreamBooth, and IP-Adapter, and validate our metrics with 32k human ratings, yielding in-depth insights into architecture and data design.

多模态生成评估基准图像生成提示对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。