提出新评估方法TIT-Score,更准确衡量长提示文本生成图像的一致性。
TIT-Score: Evaluating Long-Prompt Based Text-to-Image Alignment via Text-to-Image-to-Text Consistency
- 基于文本-图像-文本一致性设计零样本评估框架
- 在200个超长提示上测试,13模型生成2600张图并人工标注
- 相较现有方法提升7.31%判断准确率,适合研究生成一致性者
随着大规模多模态模型的快速发展,当前文本到图像(T2I)模型在短提示下已能生成高质量图像并具有良好对齐性。然而,在长而详细的提示下仍存在理解不足与生成不一致的问题。为此,我们构建了LPG-Bench——一个针对长提示文本生成的综合性评测基准,包含200个精心设计的提示,平均长度超过250词,接近多个主流商业模型的输入上限。基于该基准,我们从13个顶尖模型生成2600张图像,并进行全面的人工标注。实验发现,当前主流的T2I对齐评估指标在长提示任务中与人类偏好一致性较差。为此,我们提出一种基于文本-图像-文本一致性的新型零样本评估方法TIT,核心思想是通过比较原始提示与多模态大模型对生成图像的描述之间的一致性来量化对齐程度。TIT包括基于分数的TIT-Score和基于大语言模型的TIT-Score-LLM两种实现方式。大量实验证明,该框架在与人类判断对齐方面优于CLIP-score、LMM-score等方法,其中TIT-Score-LLM在成对准确性上相比最强基线绝对提升7.31%。LPG-Bench与TIT方法共同为评估和推动T2I模型发展提供更深入视角。所有资源将公开共享。
原文摘要 · Abstract (English)
With the rapid advancement of large multimodal models (LMMs), recent text-to-image (T2I) models can generate high-quality images and demonstrate great alignment to short prompts. However, they still struggle to effectively understand and follow long and detailed prompts, displaying inconsistent generation. To address this challenge, we introduce LPG-Bench, a comprehensive benchmark for evaluating long-prompt-based text-to-image generation. LPG-Bench features 200 meticulously crafted prompts with an average length of over 250 words, approaching the input capacity of several leading commercial models. Using these prompts, we generate 2,600 images from 13 state-of-the-art models and further perform comprehensive human-ranked annotations. Based on LPG-Bench, we observe that state-of-the-art T2I alignment evaluation metrics exhibit poor consistency with human preferences on long-prompt-based image generation. To address the gap, we introduce a novel zero-shot metric based on text-to-image-to-text consistency, termed TIT, for evaluating long-prompt-generated images. The core concept of TIT is to quantify T2I alignment by directly comparing the consistency between the raw prompt and the LMM-produced description on the generated image, which includes an efficient score-based instantiation TIT-Score and a large-language-model (LLM) based instantiation TIT-Score-LLM. Extensive experiments demonstrate that our framework achieves superior alignment with human judgment compared to CLIP-score, LMM-score, etc., with TIT-Score-LLM attaining a 7.31% absolute improvement in pairwise accuracy over the strongest baseline. LPG-Bench and TIT methods together offer a deeper perspective to benchmark and foster the development of T2I models. All resources will be made publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。