arXiv:2505.04718cs.CVcs.LG2025-05ICCV被引 6

用开源模型生成自然场景布局,支持开放式文本控制。

Lay-Your-Scene: Natural Scene Layout Generation with Diffusion Transformers

  • 用轻量开源语言模型提取文本中的场景元素
  • 提出感知长宽比的扩散Transformer,实现开放词汇布局生成
  • 可应用于图像编辑与粗略初始化,适合可控图像生成研究

我们提出Lay-Your-Scene(简称LayouSyn),一种面向自然场景的文本到布局生成新方法。以往场景布局生成方法或为封闭词汇,或依赖专有大语言模型进行开放词汇生成,限制了建模能力与可控图像生成的广泛应用。本文采用轻量级开源语言模型从文本提示中获取场景元素,并设计一种新型的、以开放词汇方式训练的长宽比感知扩散Transformer架构,实现条件布局生成。大量实验表明,LayouSyn优于现有方法,在具有挑战性的空间与数值推理基准上达到当前最优性能。此外,我们展示了两个应用:一是将大语言模型的粗略初始化与本方法无缝结合,提升生成效果;二是构建物体添加到图像的流水线,展示其在图像编辑中的潜力。

原文摘要 · Abstract (English)

We present Lay-Your-Scene (shorthand LayouSyn), a novel text-to-layout generation pipeline for natural scenes. Prior scene layout generation methods are either closed-vocabulary or use proprietary large language models for open-vocabulary generation, limiting their modeling capabilities and broader applicability in controllable image generation. In this work, we propose to use lightweight open-source language models to obtain scene elements from text prompts and a novel aspect-aware diffusion Transformer architecture trained in an open-vocabulary manner for conditional layout generation. Extensive experiments demonstrate that LayouSyn outperforms existing methods and achieves state-of-the-art performance on challenging spatial and numerical reasoning benchmarks. Additionally, we present two applications of LayouSyn. First, we show that coarse initialization from large language models can be seamlessly combined with our method to achieve better results. Second, we present a pipeline for adding objects to images, demonstrating the potential of LayouSyn in image editing applications.

场景生成扩散模型文本生成图像编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。