用户手绘即可控制生成图像,无需训练
DiffBrush:Just Painting the Art by Your Hands
- 通过调整扩散模型内部表示实现手绘引导
- 可精准控制颜色、语义和实例位置
- 适合希望快速创作的设计师与艺术家
近年来图像生成与编辑算法的快速发展使普通用户也能生成逼真图像。然而,当前AI绘画系统主要依赖文本驱动的扩散模型(T2I),难以准确捕捉用户需求,且与其他模态兼容需高昂训练成本。为此,我们提出DiffBrush,兼容T2I模型,支持用户手绘与编辑图像。通过操控和适配扩散模型的内部表示,DiffBrush在不进行额外训练的前提下,引导生成图像向用户手绘草图收敛。该方法在去噪过程中持续引导潜在空间与实例级注意力图,实现对图像颜色、语义和实例的精准控制。此外,我们提出潜变量重生成机制,优化扩散模型中随机采样的噪声,获得更合理的生成布局。最终,用户仅需在画布上粗略绘制实例遮罩(可接受颜色),DiffBrush即可自然生成对应实例于指定位置。
原文摘要 · Abstract (English)
The rapid development of image generation and editing algorithms in recent years has enabled ordinary user to produce realistic images. However, the current AI painting ecosystem predominantly relies on text-driven diffusion models (T2I), which pose challenges in accurately capturing user requirements. Furthermore, achieving compatibility with other modalities incurs substantial training costs. To this end, we introduce DiffBrush, which is compatible with T2I models and allows users to draw and edit images. By manipulating and adapting the internal representation of the diffusion model, DiffBrush guides the model-generated images to converge towards the user's hand-drawn sketches for user's specific needs without additional training. DiffBrush achieves control over the color, semantic, and instance of objects in images by continuously guiding the latent and instance-level attention map during the denoising process of the diffusion model. Besides, we propose a latent regeneration, which refines the randomly sampled noise in the diffusion model, obtaining a better image generation layout. Finally, users only need to roughly draw the mask of the instance (acceptable colors) on the canvas, DiffBrush can naturally generate the corresponding instance at the corresponding location.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。