arXiv:2605.11311cs.LGcs.CV2026-05被引 3

通过设计噪声间的关联性,提升扩散模型生成图像的多样性。

Couple to Control: Joint Initial Noise Design in Diffusion Models

论文配图:Couple to Control: Joint Initial Noise Design in Diffusion Models
图 1 · 摘自论文原文
  • 让初始噪声不再独立,而是按设计构建依赖关系。
  • 使用排斥型噪声耦合,显著提升图像集多样性且不增加采样成本。
  • 适用于需要多样背景或优化初始化的场景,兼容现有模型。

扩散模型通常从相互独立的高斯初始噪声生成图像批次。本文指出,这种独立性仅为一种可选项,实际上可设计噪声间的联合分布:每个噪声仍保持边缘标准高斯分布,使预训练模型输入分布不变,但样本间的依赖结构可由设计决定。这将初始噪声控制从单样本种子选择转变为多样本画廊的依赖结构设计。该框架涵盖多种已有方法作为特例,并自然引出新型耦合噪声构造。耦合噪声本身即可提升生成效果,无需额外采样开销;在有额外计算资源时,还可作为优化流水线的结构化初始化。实验表明,排斥型高斯耦合在SD1.5、SDXL和SD3上均提升了画廊多样性,同时基本保持提示词对齐与图像质量,且在相同采样成本下优于近期测试时噪声优化基线。子空间耦合还支持固定主体背景生成,相比专用修复基线,能生成更自然多样的背景,并可在前景保真度上灵活权衡。

原文摘要 · Abstract (English)

Diffusion models typically generate image batches from independent Gaussian initial noises. We argue that this independence assumption is only one choice within a broader class of valid joint noise designs. Instead, one can specify a coupling of the initial noises: each noise remains marginally standard Gaussian, so the pretrained diffusion model receives the same single-sample input distribution, while the dependence across samples is chosen by design. This reframes initial-noise control from selecting or optimizing individual seeds to designing the dependence structure of a multi-sample gallery. This view gives a general framework for initial-noise design, covering several existing methods as special cases and leading naturally to new coupled-noise constructions. Coupled noise can improve generation on its own without adding sampling cost, and it is flexible enough to serve as a structured initialization for optimization-based pipelines when additional computation is available. Empirically, repulsive Gaussian coupling improves gallery diversity on SD1.5, SDXL, and SD3 while largely preserving prompt alignment and image quality. It matches or outperforms recent test-time noise-optimization baselines on several diversity metrics at the same sampling cost as independent generation. Subspace couplings also support fixed-object background generation, producing diverse, natural backgrounds compared with specialized inpainting baselines, with a tunable trade-off in foreground fidelity.

扩散模型图像生成噪声设计多样性提升

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。