arXiv:2503.09419cs.CV2025-03CVPR被引 20

让扩散模型生成更稳定,抗扰动能力更强。

Alias-Free Latent Diffusion Models: Improving Fractional Shift Equivariance of Diffusion Latent Space

  • 重构注意力模块实现平移等变性,提升潜空间一致性
  • 新提出的等变损失有效抑制特征频带宽度,减少混叠
  • 适合对生成稳定性要求高的图像视频编辑场景

潜空间扩散模型(LDMs)存在生成过程不稳定的问题,微小的输入噪声扰动或偏移即可导致输出显著差异,限制了其在需要一致结果的应用中的使用。本文通过重构LDM以增强其平移等变性来改善这一问题。尽管引入抗混叠操作可部分提升等变性,但因以下挑战仍存在显著混叠和不一致:1)在变分自编码器(VAE)训练及多次U-Net推理过程中出现混叠放大;2)自注意力模块本身缺乏平移等变性。为此,本文设计了具有平移等变性的注意力模块,并提出一种等变性损失,有效抑制连续域中特征的频率带宽。所提出的无混叠扩散模型(AF-LDM)实现了强平移等变性,且对不规则形变也具备鲁棒性。大量实验表明,相比原始LDM,AF-LDM在视频编辑、图像到图像转换等多种应用中均能生成显著更一致的结果。

原文摘要 · Abstract (English)

Latent Diffusion Models (LDMs) are known to have an unstable generation process, where even small perturbations or shifts in the input noise can lead to significantly different outputs. This hinders their applicability in applications requiring consistent results. In this work, we redesign LDMs to enhance consistency by making them shift-equivariant. While introducing anti-aliasing operations can partially improve shift-equivariance, significant aliasing and inconsistency persist due to the unique challenges in LDMs, including 1) aliasing amplification during VAE training and multiple U-Net inferences, and 2) self-attention modules that inherently lack shift-equivariance. To address these issues, we redesign the attention modules to be shift-equivariant and propose an equivariance loss that effectively suppresses the frequency bandwidth of the features in the continuous domain. The resulting alias-free LDM (AF-LDM) achieves strong shift-equivariance and is also robust to irregular warping. Extensive experiments demonstrate that AF-LDM produces significantly more consistent results than vanilla LDM across various applications, including video editing and image-to-image translation.

扩散模型生成稳定等变性图像编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。