提出新评估框架,系统评测文本到图像生成模型的语义与布局一致性。
Benchmarking Layout-Guided Diffusion Models through Unified Semantic-Spatial Evaluation in Closed and Open Settings

- 构建封闭集与开放集双基准,覆盖不同复杂度提示与布局
- 31.9万张图像大规评估,给出统一语义-空间评分排名
- 揭示当前模型在复杂提示下的布局缺陷,适合研究生成可控性者
评估基于布局引导的文本到图像生成模型需同时考察语义与布局的准确性。现有基准因需精细标注而规模有限,难以全面评估。本文提出封闭集基准(C-Bench)和开放集基准(O-Bench),分别在受控与真实场景下评估模型。我们设计统一评估协议,将语义与空间准确率合并为单一得分,确保模型排名一致。基于该框架,对6个前沿扩散模型进行大规模评估,共生成并评估319,086张图像。建立综合性能排名,并提供文本与布局对齐的详细分析,揭示模型在复杂提示下的优劣。代码已开源。
原文摘要 · Abstract (English)
Evaluating layout-guided text-to-image generative models requires assessing both semantic alignment with textual prompts and spatial fidelity to prescribed layouts. Assessing layout alignment requires collecting fine-grained annotations, which is costly and labor-intensive. Consequently, current benchmarks rarely provide comprehensive layout evaluation and often remain limited in scale or coverage, making model comparison, ranking, and interpretation difficult. In this work, we introduce a closed-set benchmark (C-Bench) designed to isolate key generative capabilities while providing varying levels of complexity in both prompt structure and layout. To complement this controlled setting, we propose an open-set benchmark (O-Bench) that evaluates models using real-world prompts and layouts, offering a measure of semantic and spatial alignment in the wild. We further develop a unified evaluation protocol that combines semantic and spatial accuracy into a single score, ensuring consistent model ranking. Using our benchmarks, we conduct a large-scale evaluation of six state-of-the-art layout-guided diffusion models, totaling 319,086 generated and evaluated images. We establish a model ranking based on their overall performance and provide detailed breakdowns for text and layout alignment to enhance interpretability. Fine-grained analyses across scenarios and prompt complexities highlight the strengths and limitations of current models. Code is available at https://github.com/lparolari/cobench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。