arXiv:2506.06826cs.CVcs.AI2025-06

让多图生成背景一致,同时保持主体自由变化。

Controllable Coupled Image Generation via Diffusion Models

  • 在注意力层分离背景与主体成分,动态调节权重控制耦合程度。
  • 实验证明背景一致性、图文对齐与图像质量均优于现有方法。
  • 适合需要多图风格统一的场景,如产品展示、广告设计。

我们提出一种基于注意力层级的可控耦合图像生成方法,旨在生成多个具有相同或高度相似背景的图像。尽管背景保持一致,各图像中的中心物体仍可根据不同文本提示灵活变化。该方法通过解耦模型交叉注意力模块中的背景与主体成分,并引入随采样时间步变化的权重控制参数。通过联合优化目标(评估背景耦合度、图文对齐性及整体视觉质量),实现了更优性能。实验表明,该方法在多项指标上均优于现有方法。

原文摘要 · Abstract (English)

We provide an attention-level control method for the task of coupled image generation, where "coupled" means that multiple simultaneously generated images are expected to have the same or very similar backgrounds. While backgrounds coupled, the centered objects in the generated images are still expected to enjoy the flexibility raised from different text prompts. The proposed method disentangles the background and entity components in the model's cross-attention modules, attached with a sequence of time-varying weight control parameters depending on the time step of sampling. We optimize this sequence of weight control parameters with a combined objective that assesses how coupled the backgrounds are as well as text-to-image alignment and overall visual quality. Empirical results demonstrate that our method outperforms existing approaches across these criteria.

图像生成扩散模型可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。