arXiv:2608.19000cs.CV2026-08

让布局在扩散模型中自然浮现,实现精准且可编辑的协同设计。

Mise-en-Scène: Implicit Layout Emergence in Diffusion Transformers for Human-AI Design Co-Creation

论文配图:Mise-en-Scène: Implicit Layout Emergence in Diffusion Transformers for Human-AI Design Co-Creation
图 1 · 摘自论文原文
  • 用微调的扩散变换器隐式生成布局与图像,统一空间规划与视觉合成。
  • 在PrismLayersPlus上质量接近真实图,优于基于LLM和专用布局模型的方法。
  • 无需复杂条件机制,适配设计师对可编辑、高保真设计的需求。

从用户提供的元素自动合成图形设计,需兼顾整体构图一致性和每个元素的精确保留。现有方法通过语言模型显式预测边界框坐标,再将元素贴入,导致空间规划与视觉合成分离,常产生僵硬、比例失真的作品。我们提出一种新思路:布局能否在预训练图像编辑扩散变换器内部隐式涌现?Mise-en-Scène采用两阶段框架:第一阶段,使用小规模淘汰选择的LoRA微调扩散变换器,生成完整设计,元素排列与画布渲染同时涌现;第二阶段,通过确定性匹配放置步骤,将原始高分辨率元素精准移至草图位置,确保元素完全保真,并生成可编辑的分层设计,供设计师持续优化,而非静态图像。值得注意的是,仅需对预训练模型进行极简适应即可,无需通常用于多元素生成的专用条件机制。在大规模PrismLayersPlus基准上,Mise-en-Scène生成的设计在感知质量上最接近真实图,显著优于基于LLM的布局规划器和专用布局变换器,而匹配放置阶段进一步弥合了与真实复合图之间的保真度差距。

原文摘要 · Abstract (English)

Automating graphic design synthesis from user-provided elements requires both a coherent overall composition and the exact preservation of each asset. Existing methods predict a layout as explicit bounding-box coordinates with a language model and then paste the assets into it, which separates spatial planning from visual synthesis and tends to produce rigid, mis-scaled compositions. We instead ask whether the layout can emerge implicitly inside a pretrained image-editing diffusion transformer. We present Mise-en-Scène, a two-stage framework. In the first stage, a diffusion transformer adapted with a small, knockout-selected LoRA drafts a complete design in which the arrangement of the elements emerges jointly with the rendered canvas. In the second stage, a deterministic match-and-place step moves the original high-resolution assets to the drafted positions, which guarantees exact asset fidelity and yields an editable, layered design that a designer can keep refining rather than a flat image. Notably, a minimal adaptation of the pretrained transformer already suffices, without the specialized conditioning machinery commonly introduced for multi-element generation. On the large-scale PrismLayersPlus benchmark, the designs produced by Mise-en-Scène are the closest to the ground truth in perceived quality among all compared methods, by a wide margin over both an LLM layout planner and a specialized layout transformer, while our match-and-place stage bridges the remaining fidelity gap to the ground-truth composites.

扩散模型人机协同布局生成图像合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。