arXiv:2606.20709cs.CV2026-06被引 2

TeleStyle V2实现任意风格内容的精准迁移,支持真实与艺术风格自由组合。

TeleStyle V2: Beyond Content-Preserving Style Transfer with Self-Distillation and Distribution-Matching-Distillation

论文配图:TeleStyle V2: Beyond Content-Preserving Style Transfer with Self-Distillation and Distribution-Matching-Distillation
图 1 · 摘自论文原文
  • 通过自蒸馏构建多类型参考对,扩展风格迁移适用场景。
  • 在四种参考组合下表现稳定,优于或媲美商用模型。
  • 无需额外训练即可生成提示,适合图像编辑与创意设计应用。

给定内容参考和风格参考,内容保持型风格迁移要求模型生成既保留内容又融合风格的图像。此前的TeleStyle V1仅在真实内容与艺术风格的组合上表现良好,难以处理艺术内容与真实风格等情形。本文提出自蒸馏数据合成策略,从TeleStyle V1中构建包含真实-真实(RnR)、真实-艺术(RnS)、艺术-真实(SnR)、艺术-艺术(SnS)的三元组数据。基于这些自蒸馏数据训练的TeleStyle V2可支持所有四类参考组合。同时发现分布匹配蒸馏(Distribution-Matching Distillation)能保留基础模型的文本引导编辑能力,并缓解微调导致的内容一致性下降问题。定量评估显示,TeleStyleV2-QIE-2509-DMD性能至少不弱于Qwen-Image-Edit-2509-DMD,展现强大泛化图像编辑能力。此外,识别出原模型中参考顺序混淆问题,引入提示增强器解决。TeleStyle V2采用Qwen-Image-Edit的视觉语言编码器(Qwen2.5-VL-7B),免费生成内容与风格提示。其风格迁移效果可媲美当前领先商业模型gemini-3-pro-image-preview。

原文摘要 · Abstract (English)

Given a content reference and a style reference, content-preserving style transfer requires the model to generate stylized outputs with content and style consistency. We introduced TeleStyle V1 to tackle this problem. However, TeleStyle V1 is trained with photorealistic content reference and artistic style reference, which makes it incapable to cope with artistic content reference and realistic style reference in most cases. In this paper, we designed a Self-Distillation data synthesis strategy to construct such triplets from TeleStyle V1. Trained with such self-distilled triplets, our TeleStyle V2 supports Content-Style references in the forms of Realistic-and-Realistic (RnR), Realistic-and-Stylized (RnS), Stylized-and-Realistic (SnR), Stylized-and-Stylized (SnS). In addition, we found Distribution Matching Distillation could preserve the general text-guided image editing capability of the foundation model and fix the content consistency degradation caused by SFT process. Through quantitative evaluations, our TeleStyleV2-QIE-2509-DMD performs at least on par with Qwen-Image-Edit-2509-DMD, demonstrating strong general image editing skills beyond content-preserving style transfer. We observed the content/style reference order confusion problem in TeleStyle V1 and further introduced prompt enhancer to solve it. TeleStyle V2 uses Qwen-Image-Edit's VLM encoder, Qwen2.5-VL-7B, to generate content prompt and style prompt for free. TeleStyle V2 could achieve comparable style transfer performance with state-of-the-art commercial model, gemini-3-pro-image-preview.

风格迁移自蒸馏图像编辑多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。