arXiv:2505.00759cs.CVcs.AI2025-05被引 4

用多模态大模型做图像生成评估,更准且只需极少提示词。

Multi-Modal Language Models as Text-to-Image Model Evaluators

  • 用多模态大模型动态生成提示词,交互式评估图像生成质量。
  • 仅用1/80提示词就达到原有基准的排名效果,效率大幅提升。
  • 评估结果与人类判断相关性更高,适合快速测试生成模型性能。

文本到图像(T2I)生成模型持续进步,导致依赖静态数据集的自动评估基准逐渐过时。本文探索多模态大语言模型(MLLM)作为评估代理的潜力,通过与T2I模型交互,评估提示词生成一致性与图像美学。提出多模态文本到图像评估框架MT2IE,通过迭代生成提示词、评分生成图像,并将现有基准的评估结果与少量提示词匹配。实验表明,MT2IE在提示词生成一致性上的评分与人类判断的相关性高于已有方法;其生成的提示词能高效探测模型性能,在仅使用原有基准1/80提示词的情况下,得到相同的相对模型排名。

原文摘要 · Abstract (English)

The steady improvements of text-to-image (T2I) generative models lead to slow deprecation of automatic evaluation benchmarks that rely on static datasets, motivating researchers to seek alternative ways to evaluate the T2I progress. In this paper, we explore the potential of multi-modal large language models (MLLMs) as evaluator agents that interact with a T2I model, with the objective of assessing prompt-generation consistency and image aesthetics. We present Multimodal Text-to-Image Eval (MT2IE), an evaluation framework that iteratively generates prompts for evaluation, scores generated images and matches T2I evaluation of existing benchmarks with a fraction of the prompts used in existing static benchmarks. Moreover, we show that MT2IE's prompt-generation consistency scores have higher correlation with human judgment than scores previously introduced in the literature. MT2IE generates prompts that are efficient at probing T2I model performance, producing the same relative T2I model rankings as existing benchmarks while using only 1/80th the number of prompts for evaluation.

图像生成多模态评估框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。