arXiv:2502.02225cs.CVcs.AI2025-02被引 6

用SVD直接分析扩散模型隐空间,实现无需数据的精准图像编辑

Exploring the latent space of diffusion models directly through singular value decomposition

  • 通过奇异值分解直接解析隐空间结构
  • 仅需一对文本引导的隐码即可学习任意属性
  • 适合希望无监督控制生成结果的研究者

尽管扩散模型在生成高保真图像方面取得突破性成功,其隐空间仍相对未被充分探索。复杂的去噪轨迹和高维隐空间使得解释极具挑战。现有方法多聚焦于扩散模型中U-Net的特征空间,而非隐空间本身。本文通过奇异值分解(SVD)直接研究隐空间,发现三个可用来控制生成结果的有用性质,且无需数据收集、能保持生成图像的身份一致性。基于这些性质,我们提出一种新型图像编辑框架,仅需一对由文本提示生成的隐码即可学习任意属性。大量实验验证了该方法在图像编辑中的有效性与灵活性。代码将尽快公开,以促进该领域进一步研究与应用。

原文摘要 · Abstract (English)

Despite the groundbreaking success of diffusion models in generating high-fidelity images, their latent space remains relatively under-explored, even though it holds significant promise for enabling versatile and interpretable image editing capabilities. The complicated denoising trajectory and high dimensionality of the latent space make it extremely challenging to interpret. Existing methods mainly explore the feature space of U-Net in Diffusion Models (DMs) instead of the latent space itself. In contrast, we directly investigate the latent space via Singular Value Decomposition (SVD) and discover three useful properties that can be used to control generation results without the requirements of data collection and maintain identity fidelity generated images. Based on these properties, we propose a novel image editing framework that is capable of learning arbitrary attributes from one pair of latent codes destined by text prompts in Stable Diffusion Models. To validate our approach, extensive experiments are conducted to demonstrate its effectiveness and flexibility in image editing. We will release our codes soon to foster further research and applications in this area.

隐空间分析扩散模型图像编辑SVD

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。