QwenStyle实现高保真风格迁移,内容不变形且支持海量风格。
QwenStyle: Content-Preserving Style Transfer with Qwen-Image-Edit
- 基于Qwen-Image-Edit训练,分离内容与风格特征
- 在三类核心指标上达当前最优,保持内容一致性
- 适合需要精准内容保留的创意设计场景
内容保持型风格迁移在扩散变换器(DiTs)中因内容与风格特征内嵌而面临挑战。本文提出首个基于Qwen-Image-Edit训练的内容保持风格迁移模型QwenStyle,激活其强大的内容保留与风格定制能力。我们收集并筛选了有限特定风格的高质量数据,合成包含数千类别风格图像的三元组数据集。引入课程持续学习框架,训练QwenStyle以处理混合的干净与噪声三元组,使其在未见风格上仍保持精确的内容保留能力。QwenStyle V1在三个核心指标——风格相似性、内容一致性和美学质量——上达到当前最优表现。
原文摘要 · Abstract (English)
Content-Preserving Style transfer, given content and style references, remains challenging for Diffusion Transformers (DiTs) due to its internal entangled content and style features. In this technical report, we propose the first content-preserving style transfer model trained on Qwen-Image-Edit, which activates Qwen-Image-Edit's strong content preservation and style customization capability. We collected and filtered high quality data of limited specific styles and synthesized triplets with thousands categories of style images in-the-wild. We introduce the Curriculum Continual Learning framework to train QwenStyle with such mixture of clean and noisy triplets, which enables QwenStyle to generalize to unseen styles without degradation of the precise content preservation capability. Our QwenStyle V1 achieves state-of-the-art performance in three core metrics: style similarity, content consistency, and aesthetic quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。