用初始噪声生成布局,让多主体图像更准确不混淆。
Be Decisive: Noise-Induced Layouts for Multi-Subject Generation
- 从初始噪声中提取并动态优化布局,避免外部布局冲突。
- 在多个数据集上提升多主体生成的准确性与稳定性。
- 适合需要精确控制多主体分布的图像生成场景。
现有文本到图像扩散模型在生成多个不同主体时仍面临挑战,复杂提示常导致主体信息泄露,造成数量、属性和视觉特征的不准确。防止主体间泄露需掌握各主体的空间位置,现有方法通过外部布局控制实现,但强制预设布局常与模型由初始噪声决定的固有布局冲突,破坏模型先验。本文提出一种新方法:从初始噪声中预测与提示对齐的空间布局,并在去噪过程中持续优化。该噪声驱动布局避免了外部干预的矛盾,更好保留模型先验。方法采用小型神经网络在每个去噪步骤预测并修正布局,确保主体间边界清晰且一致。实验表明,该噪声对齐策略在文本-图像对齐度和多主体生成稳定性上优于现有布局引导方法,同时保持模型原始分布的丰富多样性。
原文摘要 · Abstract (English)
Generating multiple distinct subjects remains a challenge for existing text-to-image diffusion models. Complex prompts often lead to subject leakage, causing inaccuracies in quantities, attributes, and visual features. Preventing leakage among subjects necessitates knowledge of each subject's spatial location. Recent methods provide these spatial locations via an external layout control. However, enforcing such a prescribed layout often conflicts with the innate layout dictated by the sampled initial noise, leading to misalignment with the model's prior. In this work, we introduce a new approach that predicts a spatial layout aligned with the prompt, derived from the initial noise, and refines it throughout the denoising process. By relying on this noise-induced layout, we avoid conflicts with externally imposed layouts and better preserve the model's prior. Our method employs a small neural network to predict and refine the evolving noise-induced layout at each denoising step, ensuring clear boundaries between subjects while maintaining consistency. Experimental results show that this noise-aligned strategy achieves improved text-image alignment and more stable multi-subject generation compared to existing layout-guided techniques, while preserving the rich diversity of the model's original distribution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。