让文字生成的多层图像更协调,解决遮挡和布局不一致问题。
DreamLayer: Simultaneous Multi-Layer Generation via Diffusion Mode
- 通过注意力机制显式建模前景与背景层的关系。
- 生成的多层图像在空间布局和遮挡关系上更连贯。
- 适合需要精细图像编辑或分层分解的研究者使用。
基于扩散模型的文本驱动图像生成近期受到广泛关注。为实现更灵活的图像操作与编辑,研究已从单图生成拓展至透明图层生成与多层组合。然而,现有方法往往未能充分探索多层结构,导致层间交互不一致,如遮挡关系、空间布局和阴影等问题。本文提出DreamLayer框架,通过显式建模透明前景层与背景层的关系,实现文本驱动的多层图像协同生成。该框架包含三个核心组件:上下文感知交叉注意力(CACA)用于全局-局部信息交换,层共享自注意力(LSSA)建立鲁棒的层间连接,信息保留调和(IRH)在潜在空间细化融合细节。通过整合完整图像上下文,利用注意力机制构建层间关联,并通过调和步骤实现无缝层融合。为推动多层生成研究,我们构建了一个高质量、多样化的多层数据集,包含40万样本。大量实验与用户研究证明,DreamLayer生成的图像层更具一致性与对齐性,具备广泛适用性,包括潜在空间图像编辑和图像到图层的分解。
原文摘要 · Abstract (English)
Text-driven image generation using diffusion models has recently gained significant attention. To enable more flexible image manipulation and editing, recent research has expanded from single image generation to transparent layer generation and multi-layer compositions. However, existing approaches often fail to provide a thorough exploration of multi-layer structures, leading to inconsistent inter-layer interactions, such as occlusion relationships, spatial layout, and shadowing. In this paper, we introduce DreamLayer, a novel framework that enables coherent text-driven generation of multiple image layers, by explicitly modeling the relationship between transparent foreground and background layers. DreamLayer incorporates three key components, i.e., Context-Aware Cross-Attention (CACA) for global-local information exchange, Layer-Shared Self-Attention (LSSA) for establishing robust inter-layer connections, and Information Retained Harmonization (IRH) for refining fusion details at the latent level. By leveraging a coherent full-image context, DreamLayer builds inter-layer connections through attention mechanisms and applies a harmonization step to achieve seamless layer fusion. To facilitate research in multi-layer generation, we construct a high-quality, diverse multi-layer dataset including 400k samples. Extensive experiments and user studies demonstrate that DreamLayer generates more coherent and well-aligned layers, with broad applicability, including latent-space image editing and image-to-layer decomposition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。