arXiv:2504.08358cs.CV2025-04ICCV被引 36

用大模型自动评估图像生成质量,更准更省力。

LMM4LMM: Benchmarking and Evaluating Large-multimodal Image Generation with LMMs

  • 用大模型作为评分器,从感知、图文对齐等多维度评估图像生成效果。
  • 构建包含50,400张图和50K问答对的评测数据集,覆盖20类精细任务。
  • 适合研究图像生成、评估方法或需要高效评测的团队使用。

大型多模态模型(LMMs)在文本到图像(T2I)生成和图像到文本(I2T)理解方面取得显著进展,但生成图像仍存在感知质量差和图文对齐不佳的问题。由于人工评估成本高且效率低,亟需一种与人类偏好一致的自动评估指标。为此,我们提出EvalMi-50K,一个全面的多模态图像生成评测数据集与基准,包含:(i) 2,100个详尽提示,覆盖20个细粒度任务维度;(ii) 大规模人类偏好标注,涵盖100,000个平均意见得分(MOS)和50,000个问答对,基于24个T2I模型生成的50,400张图像。基于EvalMi-50K,我们提出LMM4LMM,一种基于LMM的多维度评估指标,涵盖感知质量、图文对应关系及任务特定准确性。大量实验表明,LMM4LMM在EvalMi-50K上达到顶尖性能,并在其他AI生成图像评估基准上展现强泛化能力,验证了EvalMi-50K与LMM4LMM的通用性。两项资源将开源于https://github.com/IntMeGroup/LMM4LMM。

原文摘要 · Abstract (English)

Recent breakthroughs in large multimodal models (LMMs) have significantly advanced both text-to-image (T2I) generation and image-to-text (I2T) interpretation. However, many generated images still suffer from issues related to perceptual quality and text-image alignment. Given the high cost and inefficiency of manual evaluation, an automatic metric that aligns with human preferences is desirable. To this end, we present EvalMi-50K, a comprehensive dataset and benchmark for evaluating large-multimodal image generation, which features (i) comprehensive tasks, encompassing 2,100 extensive prompts across 20 fine-grained task dimensions, and (ii) large-scale human-preference annotations, including 100K mean-opinion scores (MOSs) and 50K question-answering (QA) pairs annotated on 50,400 images generated from 24 T2I models. Based on EvalMi-50K, we propose LMM4LMM, an LMM-based metric for evaluating large multimodal T2I generation from multiple dimensions including perception, text-image correspondence, and task-specific accuracy. Extensive experimental results show that LMM4LMM achieves state-of-the-art performance on EvalMi-50K, and exhibits strong generalization ability on other AI-generated image evaluation benchmark datasets, manifesting the generality of both the EvalMi-50K dataset and LMM4LMM metric. Both EvalMi-50K and LMM4LMM will be released at https://github.com/IntMeGroup/LMM4LMM.

图像生成自动评估多模态测评基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。