arXiv:2409.02374cs.CVcs.LG2024-09被引 22

发现扩散模型隐空间有低维语义子空间,实现无需训练的精准图像编辑。

Exploring Low-Dimensional Subspaces in Diffusion Models for Controllable Image Editing

  • 基于噪声水平范围内后验均值预测器的局部线性特性,定位低维语义子空间。
  • 提出LOCO Edit方法,单步无训练实现可控局部编辑,支持组合与迁移。
  • 适用于文本到图像模型,适合需要高效精准编辑的研究者使用。

扩散模型虽强大,但其语义空间理解有限,难以在无监督条件下实现精确、解耦的图像生成。本文发现,在特定噪声水平范围内,扩散模型的后验均值预测器(PMP)具有局部线性特性,且其雅可比矩阵的奇异向量位于低维语义子空间。基于此,我们提出无需训练、单步完成的无监督控制编辑方法LOCO Edit,该方法识别出具备齐次性、可迁移性、可组合性和线性的编辑方向。这些性质得益于低维语义子空间结构。方法还可扩展至文本引导编辑(T-LOCO Edit),在多个文本到图像扩散模型上验证了其有效性和效率。

原文摘要 · Abstract (English)

Recently, diffusion models have emerged as a powerful class of generative models. Despite their success, there is still limited understanding of their semantic spaces. This makes it challenging to achieve precise and disentangled image generation without additional training, especially in an unsupervised way. In this work, we improve the understanding of their semantic spaces from intriguing observations: among a certain range of noise levels, (1) the learned posterior mean predictor (PMP) in the diffusion model is locally linear, and (2) the singular vectors of its Jacobian lie in low-dimensional semantic subspaces. We provide a solid theoretical basis to justify the linearity and low-rankness in the PMP. These insights allow us to propose an unsupervised, single-step, training-free LOw-rank COntrollable image editing (LOCO Edit) method for precise local editing in diffusion models. LOCO Edit identified editing directions with nice properties: homogeneity, transferability, composability, and linearity. These properties of LOCO Edit benefit greatly from the low-dimensional semantic subspace. Our method can further be extended to unsupervised or text-supervised editing in various text-to-image diffusion models (T-LOCO Edit). Finally, extensive empirical experiments demonstrate the effectiveness and efficiency of LOCO Edit.

扩散模型图像编辑低维子空间无训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。