让扩散模型只在指定区域精细控制,其余自动生成。
Localized Control in Diffusion Models via Latent Vector Prediction
- 通过预测潜在空间初始向量实现局部条件控制
- 在多个图像区域上实现高保真度精准控制
- 适合需要局部修改的图像生成任务
扩散模型已成为文本到图像生成的主流方法,能根据文本描述生成高质量图像。然而,仅靠文本实现细节控制仍需反复试错。现有方法虽引入图像级控制(如边缘、分割、深度图),但条件作用于全图,难以实现局部控制。本文提出新方法,在用户指定区域内实现精确控制,其余部分由扩散模型根据原提示自主生成。该方法引入掩码特征和额外损失项,利用任意扩散步骤中初始潜在向量的预测,增强当前步与最终样本在潜在空间的一致性。大量实验表明,本方法可有效合成高质量、局部可控的图像。
原文摘要 · Abstract (English)
Diffusion models emerged as a leading approach in text-to-image generation, producing high-quality images from textual descriptions. However, attempting to achieve detailed control to get a desired image solely through text remains a laborious trial-and-error endeavor. Recent methods have introduced image-level controls alongside with text prompts, using prior images to extract conditional information such as edges, segmentation and depth maps. While effective, these methods apply conditions uniformly across the entire image, limiting localized control. In this paper, we propose a novel methodology to enable precise local control over user-defined regions of an image, while leaving to the diffusion model the task of autonomously generating the remaining areas according to the original prompt. Our approach introduces a new training framework that incorporates masking features and an additional loss term, which leverages the prediction of the initial latent vector at any diffusion step to enhance the correspondence between the current step and the final sample in the latent space. Extensive experiments demonstrate that our method effectively synthesizes high-quality images with controlled local conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。