arXiv:2509.03516cs.CV2025-09中稿 · ICLR被引 33

新基准测试发现文生图模型在复杂场景推理上仍严重不足。

Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?

  • 构建12维评估体系,覆盖图像组成与逻辑推理两大能力。
  • 包含1080个高密度提示和约1.35万条细粒度检查项。
  • 所有模型在隐含信息推理上表现不佳,成关键瓶颈。

文生图(T2I)生成旨在从文本提示中合成图像,这些提示同时规定了必须呈现的内容并暗示可推断的信息,因此对应两个核心能力:——“组合”与“推理”。尽管近期T2I模型在组合与推理方面取得进展,现有评估基准仍存在局限:既未能全面覆盖两类能力的内部维度,也多局限于低场景密度和简单的一对一推理。为解决此问题,我们提出 extsc{T2I-CoReBench},一个全面且复杂的评估基准,用于评测T2I模型的组合与推理能力。为确保全面性,我们将组合能力围绕场景图元素(实例、属性、关系)展开,将推理能力基于推理哲学框架(演绎、归纳、溯因),构建12维评估分类体系。为提升复杂度,受现实世界复杂性驱动,每个提示均设计为更高组合密度与更强推理强度。为实现细粒度可靠评估,每条提示均配有检查清单,包含独立的“是/否”问题以逐项评估目标元素。统计上,本基准包含1,080个挑战性提示和约13,500条检查问题。在38个当前主流T2I模型上的实验表明,其组合能力在高组合场景下仍受限,而推理能力更是严重滞后,成为关键瓶颈,所有模型均难以从提示中推断出隐含元素。

原文摘要 · Abstract (English)

Text-to-image (T2I) generation aims to synthesize images from textual prompts, which jointly specify what must be shown and imply what can be inferred, which thus correspond to two core capabilities: \textbf{\textit{composition}} and \textbf{\textit{reasoning}}. Despite recent advances of T2I models in both composition and reasoning, existing benchmarks remain limited in evaluation. They not only fail to provide comprehensive coverage across and within both capabilities, but also largely restrict evaluation to low scene density and simple one-to-one reasoning. To address these limitations, we propose \textbf{\textsc{T2I-CoReBench}}, a comprehensive and complex benchmark that evaluates both composition and reasoning capabilities of T2I models. To ensure comprehensiveness, we structure composition around scene graph elements (\textit{instance}, \textit{attribute}, and \textit{relation}) and reasoning around the philosophical framework of inference (\textit{deductive}, \textit{inductive}, and \textit{abductive}), formulating a 12-dimensional evaluation taxonomy. To increase complexity, driven by the inherent real-world complexities, we curate each prompt with higher compositional density for composition and greater reasoning intensity for reasoning. To facilitate fine-grained and reliable evaluation, we also pair each evaluation prompt with a checklist that specifies individual \textit{yes/no} questions to assess each intended element independently. In statistics, our benchmark comprises 1,080 challenging prompts and around 13,500 checklist questions. Experiments across 38 current T2I models reveal that their composition capability still remains limited in high compositional scenarios, while the reasoning capability lags even further behind as a critical bottleneck, with all models struggling to infer implicit elements from prompts.

文生图评估基准推理能力组合能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。