无需标注即可精准编辑扩散模型中的局部图像属性
Unsupervised Region-Based Image Editing of Denoising Diffusion Models
- 通过投影雅可比矩阵到低维正交子空间,实现无监督语义定位
- 在多个数据集和模型上达到当前最佳效果,人脸属性编辑超越有监督方法
- 适合需要零样本局部编辑的图像生成研究者使用
尽管扩散模型在图像生成领域取得显著成功,其潜在空间仍鲜受探索。现有语义识别方法通常依赖外部监督,如文本信息或分割掩码。本文提出一种无需额外训练的方法,在预训练扩散模型的潜在空间中识别语义属性。通过将目标语义区域的雅可比矩阵投影至与非掩码区域正交的低维子空间,该方法实现了对局部掩码区域的精确语义发现与控制,无需任何标注。我们在多个数据集及多种扩散模型架构上进行了广泛实验,均达到当前最优性能。尤其在部分人脸属性编辑任务中,表现甚至优于有监督方法,展现出在局部图像属性编辑上的卓越能力。
原文摘要 · Abstract (English)
Although diffusion models have achieved remarkable success in the field of image generation, their latent space remains under-explored. Current methods for identifying semantics within latent space often rely on external supervision, such as textual information and segmentation masks. In this paper, we propose a method to identify semantic attributes in the latent space of pre-trained diffusion models without any further training. By projecting the Jacobian of the targeted semantic region into a low-dimensional subspace which is orthogonal to the non-masked regions, our approach facilitates precise semantic discovery and control over local masked areas, eliminating the need for annotations. We conducted extensive experiments across multiple datasets and various architectures of diffusion models, achieving state-of-the-art performance. In particular, for some specific face attributes, the performance of our proposed method even surpasses that of supervised approaches, demonstrating its superior ability in editing local image properties.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。