用图像反向生成评估文生图模型,更真实反映生成能力。
Image Regeneration: Evaluating Text-to-Image Model via Generating Identical Image with Multimodal Large Language Models
- 以参考图反推生成图为评估任务,避免文本匹配偏差。
- 引入GPT4V理解图像内容,使模型更准确还原原图细节。
- 提出ImageRepainter框架,提升生成质量,适合模型开发者使用。
扩散模型重塑了图像生成领域,在学术研究与艺术创作中发挥重要作用。随着新扩散模型的涌现,文生图模型的评估变得愈发重要。现有评估指标主要依赖输入文本与生成图像的直接匹配,但跨模态信息不对称导致评估结果不可靠或不完整。为此,本文提出图像再生任务:让文生图模型根据参考图像生成相同图像。通过GPT4V将参考图像与文本输入对齐,使文生图模型能理解图像内容,评估简化为生成图与参考图的直观对比。构建了涵盖内容多样性和风格多样性的两个再生数据集,用于评估当前主流扩散模型。此外,提出ImageRepainter框架,借助多模态大模型引导的迭代生成与修正,提升生成图像的内容理解与质量。大量实验验证了该框架在评估模型生成能力上的有效性。利用多模态大模型,我们证明了强大的图文对齐能力可显著提升生成图像与参考图像的一致性。
原文摘要 · Abstract (English)
Diffusion models have revitalized the image generation domain, playing crucial roles in both academic research and artistic expression. With the emergence of new diffusion models, assessing the performance of text-to-image models has become increasingly important. Current metrics focus on directly matching the input text with the generated image, but due to cross-modal information asymmetry, this leads to unreliable or incomplete assessment results. Motivated by this, we introduce the Image Regeneration task in this study to assess text-to-image models by tasking the T2I model with generating an image according to the reference image. We use GPT4V to bridge the gap between the reference image and the text input for the T2I model, allowing T2I models to understand image content. This evaluation process is simplified as comparisons between the generated image and the reference image are straightforward. Two regeneration datasets spanning content-diverse and style-diverse evaluation dataset are introduced to evaluate the leading diffusion models currently available. Additionally, we present ImageRepainter framework to enhance the quality of generated images by improving content comprehension via MLLM guided iterative generation and revision. Our comprehensive experiments have showcased the effectiveness of this framework in assessing the generative capabilities of models. By leveraging MLLM, we have demonstrated that a robust T2M can produce images more closely resembling the reference image.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。