PSDiffusion统一生成多层图像,让各层布局与视觉效果更真实
PSDiffusion: Harmonized Multi-Layer Image Generation via Layout and Appearance Alignment
- 通过全局层交互机制,协同生成多层图像
- 在基准数据集上实现更高结构合理性和视觉保真度
- 适合数字艺术、设计等领域需要高质量分层图像的场景
透明图层生成在数字艺术与设计流程中具有重要意义。现有方法通常从单张RGB图像分解出透明层,或逐层生成。尽管取得一定成果,但受限于层间共享全局上下文不足,难以建模整体布局、物理上合理的交互关系以及阴影、反光等视觉效果,且透明度质量不高。为此,我们提出PSDiffusion,一种基于预训练图像扩散模型图像构成先验的统一扩散框架,实现文本到多层图像的同步生成。具体地,引入全局层交互机制,协同生成分层图像,确保各层自身质量及层间空间与视觉关系的一致性。在基准数据集上进行大量实验表明,PSDiffusion能有效生成结构合理、视觉保真度更高的多层图像。
原文摘要 · Abstract (English)
Transparent image layer generation plays a significant role in digital art and design workflows. Existing methods typically decompose transparent layers from a single RGB image using a set of tools or generate multiple transparent layers sequentially. Despite some promising results, these methods often limit their ability to model global layout, physically plausible interactions, and visual effects such as shadows and reflections with high alpha quality due to limited shared global context among layers. To address this issue, we propose PSDiffusion, a unified diffusion framework that leverages image composition priors from pre-trained image diffusion model for simultaneous multi-layer text-to-image generation. Specifically, our method introduces a global layer interaction mechanism to generate layered images collaboratively, ensuring both individual layer quality and coherent spatial and visual relationships across layers. We include extensive experiments on benchmark datasets to demonstrate that PSDiffusion is able to outperform existing methods in generating multi-layer images with plausible structure and enhanced visual fidelity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。