arXiv:2601.00090cs.CVcs.LG2026-01被引 8

通过优化噪声提升扩散模型生成多样性,修复模式坍缩问题。

It's Never Too Late: Noise Optimization for Collapse Recovery in Trained Diffusion Models

  • 用噪声优化方法提升生成多样性,不改变原模型结构。
  • 在相同提示下生成更多样图像,质量与原模型一致。
  • 不同频率初始化的噪声可加速优化并提升搜索效果。

当前文本到图像模型存在显著的模式坍缩现象,即相同提示生成的图像趋于相似。以往工作通过引导机制或生成大量候选样本再精炼来缓解此问题。本文提出新思路:通过噪声优化实现生成多样性。实验表明,简单噪声优化目标可在保持基线模型保真度的前提下有效缓解模式坍缩。我们还分析了噪声的频率特性,发现采用不同频率分布的初始噪声能同时提升优化效率与搜索性能。在多个数据集上的实验证明,该方法在生成质量与多样性上均优于现有方案。

原文摘要 · Abstract (English)

Contemporary text-to-image models exhibit a surprising degree of mode collapse, as can be seen when sampling several images given the same text prompt. Previous work has attempted to address this issue by steering the model using guidance mechanisms, or by generating a large pool of candidates and refining them. In this work, we take a different direction and aim for diversity in generations via noise optimization. Specifically, we show that a simple noise optimization objective can mitigate mode collapse while preserving the fidelity of the base model. We also analyze the frequency characteristics of the noise and show that alternative noise initializations with different frequency profiles can improve both optimization and search. Our experiments demonstrate that noise optimization yields superior results in terms of generation quality and diversity.

扩散模型生成多样性噪声优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。