提出SCSA注意力机制,让风格迁移更准确匹配语义区域。
SCSA: A Plug-and-Play Semantic Continuous-Sparse Attention for Arbitrary Semantic Style Transfer
- 设计语义连续稀疏注意力,分别捕捉整体风格与局部纹理。
- 在相同语义区域上实现风格一致性,提升生成图像质量。
- 可即插即用,适用于各类基于注意力的风格迁移方法。
基于注意力的任意风格迁移方法(包括基于CNN、Transformer和扩散模型的方法)发展迅速,已能生成高质量风格化图像。然而,当内容图与风格图具有相同语义时,生成图像中对应语义区域的风格与风格图不一致。我们指出其根本原因在于未能考虑局部区域与语义区域之间的关系。为此,提出一种即插即用的语义连续稀疏注意力机制(SCSA),用于任意语义风格迁移——每个查询点会关注对应语义区域内的特定关键点。具体而言,语义连续注意力确保每个查询点充分关注同一语义区域内所有连续的关键点,以反映该区域的整体风格特征;语义稀疏注意力则使每个查询点聚焦于同一语义区域内最相似的稀疏关键点,以体现该区域的具体风格纹理。通过结合两者,SCSA实现了对应语义区域整体风格对齐的同时,保留了生动的纹理细节。定性与定量实验均表明,SCSA使基于注意力的任意风格迁移方法能够生成高质量的语义一致风格化图像。
原文摘要 · Abstract (English)
Attention-based arbitrary style transfer methods, including CNN-based, Transformer-based, and Diffusion-based, have flourished and produced high-quality stylized images. However, they perform poorly on the content and style images with the same semantics, i.e., the style of the corresponding semantic region of the generated stylized image is inconsistent with that of the style image. We argue that the root cause lies in their failure to consider the relationship between local regions and semantic regions. To address this issue, we propose a plug-and-play semantic continuous-sparse attention, dubbed SCSA, for arbitrary semantic style transfer -- each query point considers certain key points in the corresponding semantic region. Specifically, semantic continuous attention ensures each query point fully attends to all the continuous key points in the same semantic region that reflect the overall style characteristics of that region; Semantic sparse attention allows each query point to focus on the most similar sparse key point in the same semantic region that exhibits the specific stylistic texture of that region. By combining the two modules, the resulting SCSA aligns the overall style of the corresponding semantic regions while transferring the vivid textures of these regions. Qualitative and quantitative results prove that SCSA enables attention-based arbitrary style transfer methods to produce high-quality semantic stylized images.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。