通过精准控制扩散模型的注意力层,实现内容与风格的更好平衡。
Conditional Balance: Improving Multi-Conditioning Trade-Offs in Image Generation

- 仅向敏感注意力层注入条件信息,实现精细风格调控。
- 显著减少过度约束导致的内容失真或风格退化。
- 适合需要高质量图像生成的艺术家和设计师使用。
在图像生成中,保持内容保真度与艺术风格之间的平衡是一个关键挑战。传统风格迁移方法和现代去噪扩散概率模型(DDPM)虽致力于达成这一目标,却常因牺牲风格、内容或两者而受限。本文分析了DDPM在维持内容与风格均衡方面的能力,提出一种新方法以识别注意力层中的敏感区域,这些区域对应不同的风格特征。通过仅将条件输入引导至这些敏感层,我们的方法实现了对风格与内容的细粒度控制,显著缓解了过度约束带来的问题。实验表明,该方法提升了现有风格化技术的效果,使生成图像在风格与内容上更加协调,整体视觉质量得到改善。
原文摘要 · Abstract (English)
Balancing content fidelity and artistic style is a pivotal challenge in image generation. While traditional style transfer methods and modern Denoising Diffusion Probabilistic Models (DDPMs) strive to achieve this balance, they often struggle to do so without sacrificing either style, content, or sometimes both. This work addresses this challenge by analyzing the ability of DDPMs to maintain content and style equilibrium. We introduce a novel method to identify sensitivities within the DDPM attention layers, identifying specific layers that correspond to different stylistic aspects. By directing conditional inputs only to these sensitive layers, our approach enables fine-grained control over style and content, significantly reducing issues arising from over-constrained inputs. Our findings demonstrate that this method enhances recent stylization techniques by better aligning style and content, ultimately improving the quality of generated visual content.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。