arXiv:2506.15563cs.CV2025-06ICML被引 5

无需训练即可精准布局生成,同时保持图像真实感。

Control and Realism: Best of Both Worlds in Layout-to-Image without Training

  • 引入非局部注意力能量函数,解决定位偏差问题。
  • 采用基于Langevin动力学的自适应更新,提升图像真实性。
  • 无需训练,适合需要高控制与真实感的应用场景。

布局到图像生成旨在精确控制物体位置与排列以生成复杂场景。现有方法利用预训练文本到图像扩散模型可在不训练的情况下实现此目标,但常面临定位不准确和图像失真等问题。针对这些缺陷,我们提出一种全新的无训练方法WinWinLay。其核心包含两项策略:非局部注意力能量函数与自适应更新机制。理论上证明,传统注意力能量函数存在固有的空间分布偏差,导致物体难以均匀贴合布局指令;为此,引入非局部注意力先验重新分配注意力得分,促进物体更准确地遵循空间约束。同时,发现标准反向传播更新会偏离预训练分布,引发分布外伪影;因此提出基于Langevin动力学的自适应更新方案,确保在尊重布局约束的同时保持域内更新。大量实验表明,WinWinLay在元素定位控制与视觉保真度方面均优于当前最先进方法。

原文摘要 · Abstract (English)

Layout-to-Image generation aims to create complex scenes with precise control over the placement and arrangement of subjects. Existing works have demonstrated that pre-trained Text-to-Image diffusion models can achieve this goal without training on any specific data; however, they often face challenges with imprecise localization and unrealistic artifacts. Focusing on these drawbacks, we propose a novel training-free method, WinWinLay. At its core, WinWinLay presents two key strategies, Non-local Attention Energy Function and Adaptive Update, that collaboratively enhance control precision and realism. On one hand, we theoretically demonstrate that the commonly used attention energy function introduces inherent spatial distribution biases, hindering objects from being uniformly aligned with layout instructions. To overcome this issue, non-local attention prior is explored to redistribute attention scores, facilitating objects to better conform to the specified spatial conditions. On the other hand, we identify that the vanilla backpropagation update rule can cause deviations from the pre-trained domain, leading to out-of-distribution artifacts. We accordingly introduce a Langevin dynamics-based adaptive update scheme as a remedy that promotes in-domain updating while respecting layout constraints. Extensive experiments demonstrate that WinWinLay excels in controlling element placement and achieving photorealistic visual fidelity, outperforming the current state-of-the-art methods.

布局生成扩散模型无训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。