用新方法提升个性化图像编辑精度,让模型更懂用户想改哪里。
DreamSteerer: Enhancing Source Image Conditioned Editability using Personalized Diffusion Models
- 设计新目标函数,增强模型对源图的编辑响应能力。
- 解决模式坍缩问题,实现高保真局部修改效果。
- 可无缝接入现有个性化生成模型,效率高适合实用。
近期文本到图像个性化方法已能在少量图像输入下教会扩散模型用户指定概念,并在新场景中复用。在此基础上,个性化编辑——即利用个性化概念修改已有图像——成为更精准的引导方式。然而,直接将个性化扩散模型与文本驱动编辑框架结合时,往往在源图编辑上表现不佳。为此,本文提出 DreamSteerer,一种可插拔的增强方法,通过创新的可编辑性驱动得分蒸馏(EDSD)目标,提升个性化扩散模型对源图的条件编辑能力。同时,识别出 EDSD 存在的模式坍缩问题,提出基于空间特征引导采样的模式转移正则化以缓解。此外,对 Delta Denoising Score 框架进行两项关键改进,实现高保真局部编辑。大量实验表明,DreamSteerer 能显著提升多个 T2I 个性化基线的编辑能力,且计算开销低。
原文摘要 · Abstract (English)
Recent text-to-image personalization methods have shown great promise in teaching a diffusion model user-specified concepts given a few images for reusing the acquired concepts in a novel context. With massive efforts being dedicated to personalized generation, a promising extension is personalized editing, namely to edit an image using personalized concepts, which can provide a more precise guidance signal than traditional textual guidance. To address this, a straightforward solution is to incorporate a personalized diffusion model with a text-driven editing framework. However, such a solution often shows unsatisfactory editability on the source image. To address this, we propose DreamSteerer, a plug-in method for augmenting existing T2I personalization methods. Specifically, we enhance the source image conditioned editability of a personalized diffusion model via a novel Editability Driven Score Distillation (EDSD) objective. Moreover, we identify a mode trapping issue with EDSD, and propose a mode shifting regularization with spatial feature guided sampling to avoid such an issue. We further employ two key modifications to the Delta Denoising Score framework that enable high-fidelity local editing with personalized concepts. Extensive experiments validate that DreamSteerer can significantly improve the editability of several T2I personalization baselines while being computationally efficient.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。