arXiv:2411.02437cs.CVcs.AI2024-11被引 4

提出新评估指标TypeScore,精准衡量文生图模型的文本保真度。

TypeScore: A Text Fidelity Metric for Text-to-Image Generative Models

  • 基于图像描述模型与文本相似性集成,量化生成文本的保真程度。
  • 在多种文本风格下区分主流文生图模型性能差异,敏感度超CLIPScore。
  • 适合研究指令遵循能力、文本渲染质量的AI研究人员使用。

文生图模型的评估仍面临挑战,尽管其整体性能不断提升。现有指标如CLIPScore适用于粗粒度评估,但难以捕捉模型性能提升带来的细微差别。本文聚焦于模型生成图像中嵌入文本的呈现质量,以此作为评估模型精细指令遵循能力的切入点。为此,我们提出名为TypeScore的新评估框架,通过额外的图像描述模型,并利用原始文本与提取文本间的集成不相似性度量,敏感地评估模型生成高保真嵌入文本的能力。该指标显示了比CLIPScore更高的分辨率,可在多种指令和文本风格下区分主流图像生成模型。研究还评估了视觉语言模型(VLMs)对风格指令的遵循程度,将风格评价与文本保真度分离。通过人工评估,定量验证了该指标的有效性。全面分析探讨了文本长度、描述模型选择及当前向人类水平逼近的进展。框架揭示了图像生成中嵌入文本指令遵循能力的现存差距。

原文摘要 · Abstract (English)

Evaluating text-to-image generative models remains a challenge, despite the remarkable progress being made in their overall performances. While existing metrics like CLIPScore work for coarse evaluations, they lack the sensitivity to distinguish finer differences as model performance rapidly improves. In this work, we focus on the text rendering aspect of these models, which provides a lens for evaluating a generative model's fine-grained instruction-following capabilities. To this end, we introduce a new evaluation framework called TypeScore to sensitively assess a model's ability to generate images with high-fidelity embedded text by following precise instructions. We argue that this text generation capability serves as a proxy for general instruction-following ability in image synthesis. TypeScore uses an additional image description model and leverages an ensemble dissimilarity measure between the original and extracted text to evaluate the fidelity of the rendered text. Our proposed metric demonstrates greater resolution than CLIPScore to differentiate popular image generation models across a range of instructions with diverse text styles. Our study also evaluates how well these vision-language models (VLMs) adhere to stylistic instructions, disentangling style evaluation from embedded-text fidelity. Through human evaluation studies, we quantitatively meta-evaluate the effectiveness of the metric. Comprehensive analysis is conducted to explore factors such as text length, captioning models, and current progress towards human parity on this task. The framework provides insights into remaining gaps in instruction-following for image generation with embedded text.

文生图评估指标指令遵循文本保真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。