无需训练,通过注意力反向损失实现图文生成的精准布局控制
Layout Control and Semantic Guidance with Attention Loss Backward for T2I Diffusion Model
- 利用注意力反向损失动态调节跨注意力图,实现无训练控制
- 显著改善对象属性错配和提示词遵循问题,提升布局一致性
- 适合需要快速部署可控生成的工业级应用,尤其适配生产环境
可控图像生成是图像生成的核心需求之一,旨在生成既具创意又符合逻辑的图像,并满足额外指定条件。在后AIGC时代,可控生成依赖扩散模型,通常通过保留特定组件或引入推理干扰来实现。本文针对两个关键挑战:1)生成过程中对象属性不匹配及提示词遵循效果差;2)可控布局完成度不足。提出一种无需训练的方法,基于注意力损失反向机制,巧妙调控跨注意力图。通过将外部条件(如提示词)合理映射到注意力图,可在无需任何训练或微调的情况下实现图像生成控制。该方法有效解决属性错配与提示词遵循问题,并引入显式布局约束,已在实际生产中取得优异应用效果,有望为该领域提供重要技术参考。
原文摘要 · Abstract (English)
Controllable image generation has always been one of the core demands in image generation, aiming to create images that are both creative and logical while satisfying additional specified conditions. In the post-AIGC era, controllable generation relies on diffusion models and is accomplished by maintaining certain components or introducing inference interferences. This paper addresses key challenges in controllable generation: 1. mismatched object attributes during generation and poor prompt-following effects; 2. inadequate completion of controllable layouts. We propose a train-free method based on attention loss backward, cleverly controlling the cross attention map. By utilizing external conditions such as prompts that can reasonably map onto the attention map, we can control image generation without any training or fine-tuning. This method addresses issues like attribute mismatch and poor prompt-following while introducing explicit layout constraints for controllable image generation. Our approach has achieved excellent practical applications in production, and we hope it can serve as an inspiring technical report in this field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。