arXiv:2410.01540cs.CVcs.AI2024-10被引 13

用保边噪声提升扩散模型生成细节能力

Edge-preserving noise for diffusion models

  • 设计混合噪声方案,动态切换保边与均匀噪声
  • 在多种任务中提升生成质量,尤其擅长结构引导生成
  • 可直接微调已有模型,适合实际系统改造

传统扩散模型通常依赖各向同性的高斯噪声,对所有区域一视同仁,忽视了对高质量生成至关重要的结构信息。本文提出一种保边扩散过程,通过融合边缘感知调度器的混合噪声方案,实现从保边到各向同性噪声的平滑过渡。该方法使模型在保持全局性能的同时,有效捕捉精细结构细节。我们在扩散和流匹配框架下评估了结构感知噪声的影响,结果表明现有各向同性模型可通过保边噪声高效微调,具备良好的实用性。在无条件生成之外,本方法在笔画到图像等结构引导任务中表现更优,显著提升鲁棒性与视觉质量,在FID、KID和CLIP-score上均取得一致改进。

原文摘要 · Abstract (English)

Classical diffusion models typically rely on isotropic Gaussian noise, treating all regions uniformly and overlooking structural information important for high-quality generation. We introduce an edge-preserving diffusion process that generalizes isotropic models via a hybrid noise scheme with an edge-aware scheduler that smoothly transitions from edge-preserving to isotropic noise. This enables the model to capture fine structural details while generally maintaining global performance. We evaluate the impact of structure-aware noise in both diffusion and flow-matching frameworks, and show that existing isotropic models can be efficiently fine-tuned with edge-preserving noise, making our framework practical for adapting pre-trained systems. Beyond unconditional generation, our method particularly shows improvements in structure-guided tasks such as stroke-to-image synthesis, improving robustness and perceptual quality, as evidenced by consistent improvements across FID, KID, and CLIP-score.

扩散模型保边噪声结构生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。