用修正流生成多样可信的场景布局,小模型实现领先效果
SLayR: Scene Layout Generation with Rectified Flow
- 基于Transformer的修正流架构,支持无约束文本到布局生成
- 在多样性与合理性上超越现有方法,参数量减少3倍以上
- 自建评估基准,兼顾量化指标与可复现的人工评测
我们提出SLayR——一种基于修正流的文本到场景布局生成新方法,能够生成高质量布局并对接现有布局到图像模型。针对当前文本到图像流水线在提示模糊、缺乏约束时难以生成多样化且合理的场景布局的问题,SLayR实现了显著提升。该模型在无约束生成任务中优于现有基线(包括大语言模型),支持从开放词汇描述生成布局。为准确评估布局生成质量,我们构建了新的基准套件,包含数值指标和精心设计的可重复人工评估流程,用于衡量生成结果的合理性和多样性。实验表明,SLayR在同时达成高多样性与高合理性方面达到新基准水平,且参数量至少减少3倍。
原文摘要 · Abstract (English)
We introduce SLayR, Scene Layout Generation with Rectified flow, a novel transformer-based model for text-to-layout generation which can then be paired with existing layout-to-image models to produce images. SLayR addresses a domain in which current text-to-image pipelines struggle: generating scene layouts that are of significant variety and plausibility, when the given prompt is ambiguous and does not provide constraints on the scene. SLayR surpasses existing baselines including LLMs in unconstrained generation, and can generate layouts from an open caption set. To accurately evaluate the layout generation, we introduce a new benchmark suite, including numerical metrics and a carefully designed repeatable human-evaluation procedure that assesses the plausibility and variety of generated images. We show that our method sets a new state of the art for achieving both at the same time, while being at least 3x times smaller in the number of parameters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。