arXiv:2603.17965cs.CV2026-03被引 3

用自然语言生成可编辑的多层设计稿,支持灵活层数和透明通道。

LaDe: Unified Multi-Layered Graphic Media Generation and Decomposition

  • 通过语言扩展+4D位置编码的扩散模型,统一生成多层设计
  • 在Crello数据集上文本到图层对齐度优于现有方法
  • 适合需要可编辑设计稿的UI/UX、广告创作场景

媒体设计层生成技术可通过自然语言提示创建完全可编辑的分层设计文档,如海报、传单和标志。现有方法要么限制输出层数固定,要么要求每层仅包含连续空间区域,导致层数随设计复杂度线性增长。我们提出LaDe(分层媒体设计),一种潜在扩散框架,可生成灵活数量的语义有意义的图层。LaDe结合三个组件:基于LLM的提示扩展器,将简短用户意图转化为结构化的逐层描述以指导生成;带有4D RoPE位置编码机制的潜在扩散Transformer,联合生成完整媒体设计及其构成的RGBA图层;以及支持完整透明通道的RGBA VAE,解码每层图像。通过在训练中对图层样本进行条件约束,该统一框架支持三项任务:文本到图像生成、文本到图层媒体设计生成,以及媒体设计分解。我们在Crello测试集上将LaDe与Qwen-Image-Layered在文本到图层和图像到图层任务上进行比较。结果显示,LaDe在文本到图层生成中表现更优,提升了文本到图层对齐度,经由GPT-4o mini和Qwen3-VL两种VLM作为裁判评估验证。

原文摘要 · Abstract (English)

Media design layer generation enables the creation of fully editable, layered design documents such as posters, flyers, and logos using only natural language prompts. Existing methods either restrict outputs to a fixed number of layers or require each layer to contain only spatially continuous regions, causing the layer count to scale linearly with design complexity. We propose LaDe (Layered Media Design), a latent diffusion framework that generates a flexible number of semantically meaningful layers. LaDe combines three components: an LLM-based prompt expander that transforms a short user intent into structured per-layer descriptions that guide the generation, a Latent Diffusion Transformer with a 4D RoPE positional encoding mechanism that jointly generates the full media design and its constituent RGBA layers, and an RGBA VAE that decodes each layer with full alpha-channel support. By conditioning on layer samples during training, our unified framework supports three tasks: text-to-image generation, text-to-layers media design generation, and media design decomposition. We compare LaDe to Qwen-Image-Layered on text-to-layers and image-to-layers tasks on the Crello test set. LaDe outperforms Qwen-Image-Layered in text-to-layers generation by improving text-to-layer alignment, as validated by two VLM-as-a-judge evaluators (GPT-4o mini and Qwen3-VL).

多层生成扩散模型可编辑设计文本到图像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。