arXiv:2412.04460cs.CV2024-12被引 21

提出分层生成新方法,让文字生成图像时能自动分离前景与背景。

LayerFusion: Harmonized Multi-Layer Text-to-Image Generation with Generative Priors

  • 通过潜在扩散模型实现前后景协同生成,打破传统顺序生成限制。
  • 在视觉连贯性和图层一致性上显著优于基线方法。
  • 适合需要灵活编辑的图形设计、动画制作等创意领域使用。

大规模扩散模型在根据文本描述生成高质量图像方面已取得显著成果,广泛应用于各类场景。然而,对具有层次结构的内容(如带透明度的前后景图像)的生成仍研究不足。此类内容生成在图形设计、动画和数字艺术等领域至关重要,因层式结构便于后续编辑与组合。本文提出一种基于潜在扩散模型(LDMs)的新图像生成流程,可生成包含透明前景层(RGBA)与背景层(RGB)的双层图像。不同于现有方法的顺序生成,本方法引入协调生成机制,使两层间实现动态交互,提升输出一致性。通过大量定性与定量实验验证,结果表明该方法在视觉连贯性、图像质量及图层一致性方面均显著优于基线方法。

原文摘要 · Abstract (English)

Large-scale diffusion models have achieved remarkable success in generating high-quality images from textual descriptions, gaining popularity across various applications. However, the generation of layered content, such as transparent images with foreground and background layers, remains an under-explored area. Layered content generation is crucial for creative workflows in fields like graphic design, animation, and digital art, where layer-based approaches are fundamental for flexible editing and composition. In this paper, we propose a novel image generation pipeline based on Latent Diffusion Models (LDMs) that generates images with two layers: a foreground layer (RGBA) with transparency information and a background layer (RGB). Unlike existing methods that generate these layers sequentially, our approach introduces a harmonized generation mechanism that enables dynamic interactions between the layers for more coherent outputs. We demonstrate the effectiveness of our method through extensive qualitative and quantitative experiments, showing significant improvements in visual coherence, image quality, and layer consistency compared to baseline methods.

图像生成扩散模型分层生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。