arXiv:2506.15033cs.CV2025-06

实现艺术家级风格迁移,支持文本控制与颜色编辑。

Break Stylistic Sophon: Are We Really Meant to Confine the Imagination in Style Transfer?

  • 用语义对齐方法精准注入风格知识,避免风格漂移。
  • 结合人类反馈数据增强,减少过拟合,提升生成质量。
  • 训练免费三重扩散机制,可同时保持内容与文本控制。

本研究提出StyleWallfacer,一种统一的训练与推理框架,解决传统风格迁移中风格漂移、控制力弱等问题。首先,基于BLIP在CLIP空间生成与风格图像语义对齐的文本描述,利用大语言模型去除风格相关描述,构建语义空隙,用于微调模型以实现高效无漂移的风格注入。其次,设计基于人类反馈的数据增强策略,将微调初期生成的高质量样本加入训练集,促进渐进式学习并显著降低过拟合。最后,构建无需训练的三重扩散流程,通过替换自注意力层中的键值对(内容→风格),实现风格注入的同时保持文本控制;引入查询保留机制,减轻对原始内容的干扰。该框架首次实现风格迁移中的图像色彩编辑,达成艺术级风格迁移效果,同时完整保留原图内容。

原文摘要 · Abstract (English)

In this pioneering study, we introduce StyleWallfacer, a groundbreaking unified training and inference framework, which not only addresses various issues encountered in the style transfer process of traditional methods but also unifies the framework for different tasks. This framework is designed to revolutionize the field by enabling artist level style transfer and text driven stylization. First, we propose a semantic-based style injection method that uses BLIP to generate text descriptions strictly aligned with the semantics of the style image in CLIP space. By leveraging a large language model to remove style-related descriptions from these descriptions, we create a semantic gap. This gap is then used to fine-tune the model, enabling efficient and drift-free injection of style knowledge. Second, we propose a data augmentation strategy based on human feedback, incorporating high-quality samples generated early in the fine-tuning process into the training set to facilitate progressive learning and significantly reduce its overfitting. Finally, we design a training-free triple diffusion process using the fine-tuned model, which manipulates the features of self-attention layers in a manner similar to the cross-attention mechanism. Specifically, in the generation process, the key and value of the content-related process are replaced with those of the style-related process to inject style while maintaining text control over the model. We also introduce query preservation to mitigate disruptions to the original content. Under such a design, we have achieved high-quality image-driven style transfer and text-driven stylization, delivering artist-level style transfer results while preserving the original image content. Moreover, we achieve image color editing during the style transfer process for the first time.

风格迁移文本控制扩散模型色彩编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。