无需训练,用注意力调节实现风格迁移中内容与风格的平衡
HAM: A Training-Free Style Transfer Approach via Heterogeneous Attention Modulation for Diffusion Models
- 通过异构注意力调制动态控制不同注意力机制
- 在多个指标上达到当前最优,保留内容细节并捕捉复杂风格
- 适合需要快速风格迁移且不希望修改模型的研究者
扩散模型在图像生成领域表现出色,尤其在风格迁移任务中。现有方法通常依赖预训练模型的特征提取能力,并通过外部控制路径显式施加风格信号,但常无法准确捕捉复杂风格参考或保持用户输入内容图像的身份信息,陷入风格-内容权衡困境。为此,我们提出一种无需训练的风格迁移方法HAM(异构注意力调制),通过初始化风格噪声并引入全局注意力调节(GAR)与局部注意力移植(LAT)机制,在图像/文本引导的风格迁移过程中有效保护内容身份信息,缓解风格-内容失衡问题。该方法在多组定性与定量实验中表现优异,各项指标达当前最优水平。
原文摘要 · Abstract (English)
Diffusion models have demonstrated remarkable performance in image generation, particularly within the domain of style transfer. Prevailing style transfer approaches typically leverage pre-trained diffusion models' robust feature extraction capabilities alongside external modular control pathways to explicitly impose style guidance signals. However, these methods often fail to capture complex style reference or retain the identity of user-provided content images, thus falling into the trap of style-content balance. Thus, we propose a training-free style transfer approach via $\textbf{h}$eterogeneous $\textbf{a}$ttention $\textbf{m}$odulation ($\textbf{HAM}$) to protect identity information during image/text-guided style reference transfer, thereby addressing the style-content trade-off challenge. Specifically, we first introduces style noise initialization to initialize latent noise for diffusion. Then, during the diffusion process, it innovatively employs HAM for different attention mechanisms, including Global Attention Regulation (GAR) and Local Attention Transplantation (LAT), which better preserving the details of the content image while capturing complex style references. Our approach is validated through a series of qualitative and quantitative experiments, achieving state-of-the-art performance on multiple quantitative metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。