arXiv:2608.29210cs.AI2026-09

新评测框架Imag-Eval揭示图像生成模型的指令理解瓶颈

Imag-Eval: a language-grounded framework for interpretable Text-to-Image instruction following evaluation

论文配图:Imag-Eval: a language-grounded framework for interpretable Text-to-Image instruction following evaluation
图 1 · 摘自论文原文
  • 分离语言复杂度与组合难度,独立控制实例数和规则绑定
  • 1140个提示+8842条规则,发现规则数量决定生成质量
  • 适合研究文本到图像对齐机制或评估模型鲁棒性的学者

文本到图像(T2I)模型虽在视觉保真度上表现优异,但现有评测基准常缺乏可解释性且诊断能力不足。传统技能评估忽略关键失败模式,如缺失部件导致的全局不连贯或物理上不可能的配置(如漂浮物体)。此外,提示难度通常仅沿单一维度控制(如提示长度或生成元素数量)。为此,我们提出Imag-Eval,一个受控基准,用于评估T2I模型将复合自然语言指令映射为视觉输出的能力。不同于以往将表面语言复杂度与组合难度混淆的做法,Imag-Eval通过独立变化实例数量与约束组合(规则),避免误差传播,实现细粒度、可解释的跨模态指令遵循失败分析。该基准包含1,140个提示和8,842个组合规则,并在多个前沿模型上进行评估。结合对超过2,000个来自同期基准提示的额外研究,结果表明:对于结构化技能,组合难度主要由实际落地的规则数量及其与实例的绑定关系决定,而非仅由提示长度决定。

原文摘要 · Abstract (English)

Text-to-Image (T2I) models have recently achieved impressive visual fidelity, yet their evaluation remains constrained by benchmarks that are often difficult to interpret and insufficiently diagnostic. Existing skill-based evaluations tend to overlook critical failure modes that strongly impact usability but fall outside standard taxonomies, such as global incoherence arising from missing parts or physically implausible configurations (e.g., floating objects). In addition, prompt difficulty is typically controlled along a single dimension; either prompt length or the number of elements to generate. To address these limitations, we introduce Imag-Eval, a controlled benchmark designed to assess how T2I models ground compositional natural-language instructions into visual outputs. Unlike prior work that conflates surface linguistic complexity with compositional difficulty, Imag-Eval explicitly seeks to disentangles these factors by independently varying both the number of instances and the combination of constraints (rules), while avoiding error propagation. This design enables fine-grained and interpretable analysis of where cross-modal instruction following fails. Our benchmark comprises 1,140 prompts and 8,842 combined rules, and we evaluate it on several state-of-the-art models. Complementing this analysis with an additional study of over 2,000 prompts from a concurrent benchmark, our results suggest that, for structured skills, compositional difficulty is primarily governed by the number of grounded rules and their binding to instances,, rather than by prompt length alone.

图像生成评测基准指令遵循可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。