arXiv:2409.14704cs.CVcs.AI2024-09EMNLP被引 1

提出新评估方法VLEU,量化文本生成图像模型的泛化能力。

VLEU: a Method for Automatic Evaluation for Generalizability of Text-to-Image Models

  • 用大语言模型采样多样文本提示,覆盖所有可能输入
  • 通过CLIP计算图文一致性,用KL散度衡量泛化性能
  • 适合评估模型在复杂文本下的生成能力,研究者必看

文本到图像(T2I)模型进展显著,但现有评估指标难以充分衡量模型处理多样化文本提示的能力,而这对其泛化性至关重要。为此,我们提出一种新指标VLEU(Visual Language Evaluation Understudy)。VLEU利用大语言模型从视觉文本域(即T2I模型的所有可能输入文本集合)中采样,生成大量多样化提示。基于CLIP模型评估这些提示生成图像与文本的对齐程度。VLEU通过计算视觉文本边际分布与模型生成图像条件分布之间的Kullback-Leibler散度,量化模型的泛化能力。该指标为不同T2I模型的比较及微调过程中的性能跟踪提供量化依据。实验表明,VLEU能有效评估多种T2I模型的泛化能力,是未来文本到图像合成研究的关键评估工具。

原文摘要 · Abstract (English)

Progress in Text-to-Image (T2I) models has significantly improved the generation of images from textual descriptions. However, existing evaluation metrics do not adequately assess the models' ability to handle a diverse range of textual prompts, which is crucial for their generalizability. To address this, we introduce a new metric called Visual Language Evaluation Understudy (VLEU). VLEU uses large language models to sample from the visual text domain, the set of all possible input texts for T2I models, to generate a wide variety of prompts. The images generated from these prompts are evaluated based on their alignment with the input text using the CLIP model.VLEU quantifies a model's generalizability by computing the Kullback-Leibler divergence between the marginal distribution of the visual text and the conditional distribution of the images generated by the model. This metric provides a quantitative way to compare different T2I models and track improvements during model finetuning. Our experiments demonstrate the effectiveness of VLEU in evaluating the generalization capability of various T2I models, positioning it as an essential metric for future research in text-to-image synthesis.

文本生成评估指标图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。