通过分层生成透明图像块,实现对场景中物体属性与位置的精细控制。
Generating Compositional Scenes via Text-to-image RGBA Instance Generation
- 将物体作为带透明度的RGBA图像分步生成,确保属性可控。
- 多层组合生成使复杂场景构建更灵活,支持精确调整物体位置和外观。
- 适合需要高精度编辑复杂图像的用户,如设计、影视特效领域。
文本到图像的扩散生成模型虽能生成高质量图像,但需繁琐的提示工程。引入布局条件可提升可控性,但现有方法缺乏布局编辑能力及对物体属性的细粒度控制。多层生成概念具有潜力,但同时生成图像实例与场景组合会限制对物体属性、三维空间相对位置及场景操作的控制。本文提出一种新型多阶段生成范式,旨在实现细粒度控制、灵活性与交互性。为确保实例属性可控,我们设计了一种新训练范式,使扩散模型能生成带有透明信息的孤立场景组件(RGBA图像)。通过预生成的实例,我们引入多层复合生成过程,实现组件在真实场景中的平滑组装。实验表明,我们的RGBA扩散模型可生成多样且高质量的实例,并精确控制物体属性。通过多层组合,本方法能够从复杂提示构建并操纵图像,对物体外观与位置的控制优于现有方法。
原文摘要 · Abstract (English)
Text-to-image diffusion generative models can generate high quality images at the cost of tedious prompt engineering. Controllability can be improved by introducing layout conditioning, however existing methods lack layout editing ability and fine-grained control over object attributes. The concept of multi-layer generation holds great potential to address these limitations, however generating image instances concurrently to scene composition limits control over fine-grained object attributes, relative positioning in 3D space and scene manipulation abilities. In this work, we propose a novel multi-stage generation paradigm that is designed for fine-grained control, flexibility and interactivity. To ensure control over instance attributes, we devise a novel training paradigm to adapt a diffusion model to generate isolated scene components as RGBA images with transparency information. To build complex images, we employ these pre-generated instances and introduce a multi-layer composite generation process that smoothly assembles components in realistic scenes. Our experiments show that our RGBA diffusion model is capable of generating diverse and high quality instances with precise control over object attributes. Through multi-layer composition, we demonstrate that our approach allows to build and manipulate images from highly complex prompts with fine-grained control over object appearance and location, granting a higher degree of control than competing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。