arXiv:2409.00313cs.CV2024-09被引 8

无需训练,用草图精准控制图像生成过程

Training-Free Sketch-Guided Diffusion with Latent Optimization

  • 通过扩散模型的交叉注意力图追踪草图关键结构
  • 在生成过程中优化潜在空间以匹配草图布局
  • 适合需要精确布局控制的创意设计场景

基于先进的扩散模型,文本到图像(T2I)生成模型已展现出生成多样化、高质量图像的能力。然而,在真实内容创作中,如何让用户对生成结果实现精确控制仍面临重大挑战。本文提出一种无需训练的创新流程,将现有文本到图像生成模型扩展为可接受草图为额外条件。为生成与输入草图布局和结构高度一致的新图像,我们发现扩散模型的交叉注意力图能有效追踪草图的核心特征。为此,我们引入潜在空间优化(latent optimization),在生成过程的每个中间步骤利用交叉注意力图对噪声潜在变量进行精调,确保生成图像严格遵循参考草图的结构。通过该方法,显著提升了图像生成的准确性,使用户在内容创作中获得更强的控制力与定制化能力。

原文摘要 · Abstract (English)

Based on recent advanced diffusion models, Text-to-image (T2I) generation models have demonstrated their capabilities to generate diverse and high-quality images. However, leveraging their potential for real-world content creation, particularly in providing users with precise control over the image generation result, poses a significant challenge. In this paper, we propose an innovative training-free pipeline that extends existing text-to-image generation models to incorporate a sketch as an additional condition. To generate new images with a layout and structure closely resembling the input sketch, we find that these core features of a sketch can be tracked with the cross-attention maps of diffusion models. We introduce latent optimization, a method that refines the noisy latent at each intermediate step of the generation process using cross-attention maps to ensure that the generated images adhere closely to the desired structure outlined in the reference sketch. Through latent optimization, our method enhances the accuracy of image generation, offering users greater control and customization options in content creation.

图像生成草图控制扩散模型无训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。