首个综合评估布局引导图像生成中语义与空间对齐的基准测试
7Bench: a Comprehensive Benchmark for Layout-guided Text-to-image Models
- 构建七类挑战场景,联合评估文本与布局对齐
- 提出布局对齐分数,量化空间准确性表现
- 适合研究合成数据、可控图像生成的开发者使用
布局引导的文本到图像模型通过显式依赖元素的空间布局,提升生成过程的控制力,已广泛应用于内容创作和合成数据生成。其核心挑战在于实现图像、文本提示与布局之间的精确对齐,确保语义一致性和空间准确性。尽管现有基准已评估文本对齐,但布局对齐仍被忽视,且缺乏同时评估两者的基准。这一空白限制了对模型空间保真度的评估,而空间错误会引入噪声,降低合成数据质量。本文提出7Bench,首个针对布局引导图像生成中语义与空间对齐的综合性基准。它包含七类挑战性场景,涵盖对象生成、色彩保真度、属性识别、物体间关系及空间控制。我们设计了一套评估协议,在现有框架基础上引入布局对齐分数以衡量空间精度。基于7Bench,我们评估多个前沿扩散模型,揭示其在不同对齐任务中的优劣。该基准已开源:https://github.com/Elizzo/7Bench。
原文摘要 · Abstract (English)
Layout-guided text-to-image models offer greater control over the generation process by explicitly conditioning image synthesis on the spatial arrangement of elements. As a result, their adoption has increased in many computer vision applications, ranging from content creation to synthetic data generation. A critical challenge is achieving precise alignment between the image, textual prompt, and layout, ensuring semantic fidelity and spatial accuracy. Although recent benchmarks assess text alignment, layout alignment remains overlooked, and no existing benchmark jointly evaluates both. This gap limits the ability to evaluate a model's spatial fidelity, which is crucial when using layout-guided generation for synthetic data, as errors can introduce noise and degrade data quality. In this work, we introduce 7Bench, the first benchmark to assess both semantic and spatial alignment in layout-guided text-to-image generation. It features text-and-layout pairs spanning seven challenging scenarios, investigating object generation, color fidelity, attribute recognition, inter-object relationships, and spatial control. We propose an evaluation protocol that builds on existing frameworks by incorporating the layout alignment score to assess spatial accuracy. Using 7Bench, we evaluate several state-of-the-art diffusion models, uncovering their respective strengths and limitations across diverse alignment tasks. The benchmark is available at https://github.com/Elizzo/7Bench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。