用扩散Transformer生成带透明通道的分层PSD文件
OmniPSD: Layered PSD Generation with Diffusion Transformer
- 通过空间注意力学习多层布局与组合关系
- 支持文本生成和图像分解,保持透明度与结构一致
- 适合需要可编辑设计稿的UI/UX设计师使用
扩散模型在图像生成与编辑方面取得显著进展,但生成或重构带有透明通道的分层PSD文件仍具挑战。本文提出OmniPSD,一个基于Flux生态的统一扩散框架,通过上下文学习实现文本到PSD生成与图像到PSD分解。在文本生成中,OmniPSD将多层按空间位置排布于单个画布,并通过空间注意力学习其组合关系,生成语义连贯且层次分明的图层。在图像分解中,采用迭代上下文编辑策略,逐步提取并擦除文本与前景内容,从单张扁平图像重建可编辑的PSD图层。引入RGBA-VAE作为辅助表示模块,在不干扰结构学习的前提下保留透明度信息。在新构建的RGBA分层数据集上的大量实验表明,OmniPSD实现了高保真生成、结构一致性与透明度感知,为分层设计的生成与分解提供了新的扩散变压器范式。
原文摘要 · Abstract (English)
Recent advances in diffusion models have greatly improved image generation and editing, yet generating or reconstructing layered PSD files with transparent alpha channels remains highly challenging. We propose OmniPSD, a unified diffusion framework built upon the Flux ecosystem that enables both text-to-PSD generation and image-to-PSD decomposition through in-context learning. For text-to-PSD generation, OmniPSD arranges multiple target layers spatially into a single canvas and learns their compositional relationships through spatial attention, producing semantically coherent and hierarchically structured layers. For image-to-PSD decomposition, it performs iterative in-context editing, progressively extracting and erasing textual and foreground components to reconstruct editable PSD layers from a single flattened image. An RGBA-VAE is employed as an auxiliary representation module to preserve transparency without affecting structure learning. Extensive experiments on our new RGBA-layered dataset demonstrate that OmniPSD achieves high-fidelity generation, structural consistency, and transparency awareness, offering a new paradigm for layered design generation and decomposition with diffusion transformers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。