arXiv:2410.09566cs.CVcs.HC2024-10ECCV

用文字控制艺术风格迁移,无需参考图且速度快。

Bridging Text and Image for Artist Style Transfer via Contrastive Learning

  • 基于对比学习对齐文本与图像风格,实现文本驱动的风格迁移。
  • 在512x512图像上仅需0.03秒,不依赖在线微调。
  • 适合需要快速、灵活生成艺术家风格图像的研究者或创作者。

图像风格迁移近年来备受关注,但通常需额外的风格参考图像,灵活性不足。使用文本描述风格是最自然的方式,尤其能表达抽象艺术风格(如特定艺术家或艺术流派)。本文提出一种基于对比学习的艺术风格迁移方法(CLAST),利用先进的图文编码器(如CLIP)提取文本中的风格描述,并通过监督对比训练实现风格与文本的一致性对齐。为此,我们设计了一种新颖高效的adaLN状态空间模型,用于探索风格与内容的融合。最终实现了纯文本驱动的图像风格迁移。大量实验表明,该方法在艺术风格迁移任务中优于现有最先进方法。更重要的是,它无需在线微调,可在0.03秒内完成512x512图像的生成。

原文摘要 · Abstract (English)

Image style transfer has attracted widespread attention in the past few years. Despite its remarkable results, it requires additional style images available as references, making it less flexible and inconvenient. Using text is the most natural way to describe the style. More importantly, text can describe implicit abstract styles, like styles of specific artists or art movements. In this paper, we propose a Contrastive Learning for Artistic Style Transfer (CLAST) that leverages advanced image-text encoders to control arbitrary style transfer. We introduce a supervised contrastive training strategy to effectively extract style descriptions from the image-text model (i.e., CLIP), which aligns stylization with the text description. To this end, we also propose a novel and efficient adaLN based state space models that explore style-content fusion. Finally, we achieve a text-driven image style transfer. Extensive experiments demonstrate that our approach outperforms the state-of-the-art methods in artistic style transfer. More importantly, it does not require online fine-tuning and can render a 512x512 image in 0.03s.

风格迁移文本控制对比学习高效生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。