arXiv:2412.08503cs.CV2024-12CVPR被引 39

让文字精准控制图像风格,避免风格错乱和内容偏差。

StyleStudio: Text-Driven Style Transfer with Selective Control of Style Elements

  • 用跨模态自适应归一化融合风格与文本特征,提升对齐度。
  • 引入风格分类无指导机制,可选择性控制风格元素。
  • 早期阶段用教师模型稳定布局,减少生成瑕疵,适合快速集成。

文本驱动的风格迁移旨在将参考图像的风格与文本提示描述的内容融合。尽管文本到图像模型的进展提升了风格转换的细腻度,但仍面临过度拟合参考风格、风格控制不足以及与文本内容不匹配等挑战。本文提出三种互补策略:首先,引入跨模态自适应实例归一化(AdaIN)机制,实现风格与文本特征更优融合,增强对齐;其次,设计基于风格的分类无指导引导(SCFG)方法,实现对风格元素的选择性控制,减少无关影响;最后,在生成初期引入教师模型以稳定空间布局,缓解伪影问题。大量实验表明,该方法显著提升了风格迁移质量与文本提示的一致性。此外,本方法可无缝集成至现有风格迁移框架,无需微调。

原文摘要 · Abstract (English)

Text-driven style transfer aims to merge the style of a reference image with content described by a text prompt. Recent advancements in text-to-image models have improved the nuance of style transformations, yet significant challenges remain, particularly with overfitting to reference styles, limiting stylistic control, and misaligning with textual content. In this paper, we propose three complementary strategies to address these issues. First, we introduce a cross-modal Adaptive Instance Normalization (AdaIN) mechanism for better integration of style and text features, enhancing alignment. Second, we develop a Style-based Classifier-Free Guidance (SCFG) approach that enables selective control over stylistic elements, reducing irrelevant influences. Finally, we incorporate a teacher model during early generation stages to stabilize spatial layouts and mitigate artifacts. Our extensive evaluations demonstrate significant improvements in style transfer quality and alignment with textual prompts. Furthermore, our approach can be integrated into existing style transfer frameworks without fine-tuning.

风格迁移文本生成可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。