通过引导势后验初始化,提升生成图像多样性。
Initialization is Half the Battle: Generating Diverse Images from a Guidance Potential Posterior

- 从引导势后验采样初始噪声,重加权先验以增强多样性。
- 在图文生成任务中,显著提升图像多样性且不损失质量。
- 兼容扩散与流匹配模型,适合追求多样性的生成应用。
尽管生成模型具有出色的保真度,但常面临模式崩溃问题。现有提升多样性的方法多集中在生成轨迹上干预。我们发现一个关键疏漏:标准高斯初始化因忽略引导势景观,常导致轨迹坍缩至主导模式。本文提出从引导势后验选择初始噪声,有效将先验重加权至多样性丰富的区域。为高效采样该分布,我们引入多样性诱导初始化(DivIn),利用朗之万动力学主动导航初始化空间,将初始噪声从坍缩区域推开并锚定于有效数据流形。本方法作为推理时多样性增强策略,兼容扩散与流匹配模型。大量实验表明,DivIn在类别到图像和文本到图像场景中均表现优异。此外,由于其与轨迹方法正交,二者结合可显著拓展多样性-质量帕累托前沿,超越单独使用任一方法的效果。
原文摘要 · Abstract (English)
Despite the remarkable fidelity of generative models, they frequently suffer from mode collapse. Existing strategies for enhancing diversity predominantly focus on intervening during the generation trajectory. We identify a critical oversight that the standard Gaussian initialization often causes trajectories to collapse into dominant modes because it is agnostic to the guidance potential landscape. In this work, we formulate selecting the initial noise from a guidance potential posterior, which effectively re-weights the prior towards diversity-rich regions. To sample from this distribution efficiently, we introduce Diversity-inducing Initialization (DivIn), which leverages Langevin dynamics to actively navigate the initialization landscape, steering initial noise away from collapsing regions while anchoring them to the valid data manifold. Our method serves as an inference-time diversity enhancement compatible with both diffusion and flow matching models. Extensive experiments show that DivIn exhibits a superior performance in both class-to-image and text-to-image scenarios. Furthermore, we highlight that as DivIn is orthogonal to trajectory-based methods, combining them significantly expands the diversity-quality Pareto frontier beyond what either achieves in isolation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。