arXiv:2601.20354cs.CV2026-01中稿 · ICLR被引 9

新基准测试文生图模型的空间理解能力,发现其仍难处理复杂空间关系。

Everything in Its Place: Benchmarking Spatial Intelligence of Text-to-Image Models

  • 设计1230个长而信息密集的提示,覆盖25个真实场景和10类空间问题。
  • 21个顶尖模型测试显示,高阶空间推理仍是主要瓶颈,性能提升有限。
  • 构建1.5万对图文数据集,微调后模型在空间关系上表现显著改善。

文生图(T2I)模型虽在生成高质量图像方面取得显著进展,但常在处理复杂空间关系(如空间感知、推理或交互)时失败。现有基准因提示过短或信息稀疏,忽视了这些关键维度。本文提出SpatialGenEval新基准,系统评估T2I模型的空间智能,包含1,230个长且信息密集的提示,覆盖25个真实场景,每个提示涵盖10类空间子领域及对应10道多选问答题,涉及物体位置、布局、遮挡与因果关系等。对21个前沿模型的广泛评估表明,高阶空间推理仍是主要瓶颈。此外,为证明信息密集设计的实用性,我们构建了SpatialT2I数据集,包含15,400对文本-图像配对,通过重写提示确保图像一致性并保留信息密度。在当前基础模型(Stable Diffusion-XL、Uniworld-V1、OmniGen2)上微调后,性能平均提升+4.2%、+5.7%、+4.4%,空间关系更逼真,验证了以数据为中心提升空间智能的新范式。

原文摘要 · Abstract (English)

Text-to-image (T2I) models have achieved remarkable success in generating high-fidelity images, but they often fail in handling complex spatial relationships, e.g., spatial perception, reasoning, or interaction. These critical aspects are largely overlooked by current benchmarks due to their short or information-sparse prompt design. In this paper, we introduce SpatialGenEval, a new benchmark designed to systematically evaluate the spatial intelligence of T2I models, covering two key aspects: (1) SpatialGenEval involves 1,230 long, information-dense prompts across 25 real-world scenes. Each prompt integrates 10 spatial sub-domains and corresponding 10 multi-choice question-answer pairs, ranging from object position and layout to occlusion and causality. Our extensive evaluation of 21 state-of-the-art models reveals that higher-order spatial reasoning remains a primary bottleneck. (2) To demonstrate that the utility of our information-dense design goes beyond simple evaluation, we also construct the SpatialT2I dataset. It contains 15,400 text-image pairs with rewritten prompts to ensure image consistency while preserving information density. Fine-tuned results on current foundation models (i.e., Stable Diffusion-XL, Uniworld-V1, OmniGen2) yield consistent performance gains (+4.2%, +5.7%, +4.4%) and more realistic effects in spatial relations, highlighting a data-centric paradigm to achieve spatial intelligence in T2I models.

文生图空间推理评测基准数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。