arXiv:2505.07481cs.CV2025-05被引 2

解决扩散模型隐空间插值退化问题,提升图像生成质量

Addressing degeneracies in latent interpolation for diffusion models

  • 提出简单归一化方案,改善多输入隐空间插值稳定性
  • 实验显示该方法显著降低退化现象,提升FID与CLIP距离指标
  • 适用于图像增强与形态变换,尤其在大数量输入时更有效

图像生成扩散模型在深度数据增强和图像形态变换中应用日益广泛。通过反转输入图像生成隐空间表示并进行插值,可生成混合特征的新图像。然而我们发现,当输入数量较多时,该插值过程易产生退化结果。本文从理论与实验两方面分析了退化成因,并提出一种简单的归一化方法作为解决方案。该方法在需要隐空间插值的场景下易于实现。通过FID和CLIP嵌入距离评估图像质量,实验表明基线插值方法在退化现象明显前已导致质量下降;而本文方法显著缓解退化问题,在非退化情形下亦能提升质量指标。

原文摘要 · Abstract (English)

There is an increasing interest in using image-generating diffusion models for deep data augmentation and image morphing. In this context, it is useful to interpolate between latents produced by inverting a set of input images, in order to generate new images representing some mixture of the inputs. We observe that such interpolation can easily lead to degenerate results when the number of inputs is large. We analyze the cause of this effect theoretically and experimentally, and suggest a suitable remedy. The suggested approach is a relatively simple normalization scheme that is easy to use whenever interpolation between latents is needed. We measure image quality using FID and CLIP embedding distance and show experimentally that baseline interpolation methods lead to a drop in quality metrics long before the degeneration issue is clearly visible. In contrast, our method significantly reduces the degeneration effect and leads to improved quality metrics also in non-degenerate situations.

扩散模型隐空间插值图像生成数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。