arXiv:2506.02161cs.CV2025-06被引 61

新基准测试评估文生图模型对复杂指令的理解能力。

TIIF-Bench: How Does Your T2I Model Follow Your Instructions?

  • 构建5000个多维度指令,分三难度等级,含长短版本保持语义一致。
  • 提出新型全局归一化编辑距离指标,精准衡量文本与图像匹配度。
  • 利用视觉语言模型自动评分,支持细粒度、可复现的模型评估。

文生图(T2I)模型的快速发展推动了AI内容生成的新阶段,其对用户指令的理解与执行能力日益增强。然而,现有评估基准在提示多样性、复杂性及评价指标精细度方面存在不足,难以准确衡量文本指令与生成图像之间的细粒度对齐性能。本文提出TIIF-Bench(Text-to-Image Instruction Following Benchmark),系统评估T2I模型对复杂指令的解析与遵循能力。该基准包含5000个按多维度组织的提示,分为三个难度层级,并为每个提示提供短版与长版(语义一致)以评估长度鲁棒性。我们提出一种新型全局归一化编辑距离(GNED)指标用于文本渲染评估,并为每个提示提供多种长宽比的参考图像以检验风格控制能力。此外,收集100个设计师级高质量提示,覆盖多样化场景。为实现可扩展、细粒度评估,探索利用大型视觉语言模型(VLMs)作为自动化二元评判器的最佳范式。通过大量消融实验,构建出可复现、可解释、可靠的评估器,能有效识别T2I模型输出的细微差异。基于TIIF-Bench对主流T2I模型的全面评测,揭示了当前系统的优劣与现有评估基准的局限性。

原文摘要 · Abstract (English)

The rapid advancements of Text-to-Image (T2I) models have ushered in a new phase of AI-generated content, marked by their growing ability to interpret and follow user instructions. However, existing T2I model evaluation benchmarks fall short in limited prompt diversity and complexity, as well as coarse evaluation metrics, making it difficult to evaluate the fine-grained alignment performance between textual instructions and generated images. In this paper, we present TIIF-Bench Text-to-Image Instruction Following Benchmark), aiming to systematically assess T2I models' ability in interpreting and following intricate textual instructions. TIIF-Bench comprises 5,000 prompts organized along multiple dimensions and categorized into three levels of difficulty and complexity. To rigorously evaluate robustness to prompt length, each prompt is provided in both short and long versions with identical core semantics. We further propose a novel Global Normalized Edit Distance (GNED) metric for text rendering and provide aspect-ratio-diverse reference images for each prompt to assess style control. In addition, we collect 100 high-quality designer-level prompts covering diverse scenarios for comprehensive evaluation. To enable scalable and fine-grained evaluation, we explore the best paradigm for leveraging the world knowledge encoded in large Vision-Language Models (VLMs) as automated binary evaluators. Through extensive ablations, we develop a fully reproducible evaluator that provides interpretable reasoning and reliable verification, enabling our benchmark to discern subtle variations in T2I model outputs. Through comprehensive benchmarking of mainstream T2I models on TIIF-Bench, we analyze the strengths and weaknesses of current T2I systems and reveal the limitations of existing evaluation benchmarks. Project Page: https://a113n-w3i.github.io/TIIF_Bench/.

文生图指令遵循评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。