arXiv:2601.20175cs.CV2026-01被引 6

TeleStyle实现图像视频的保内容风格迁移,效果更真实一致。

TeleStyle: Content-Preserving Style Transfer in Images and Videos

  • 基于Qwen-Image-Edit构建,采用课程持续学习训练
  • 在三个核心指标上达当前最优,风格相似度高、内容保真强
  • 适合需要高质量风格迁移的视觉生成应用

内容保真风格迁移在扩散变换器(DiTs)中面临挑战,因内容与风格特征在内部表示中纠缠。本文提出TeleStyle,一个轻量高效且适用于图像和视频的风格迁移模型。基于Qwen-Image-Edit,利用其出色的保内容与风格定制能力。为支持有效训练,我们构建了高质量特定风格数据集,并合成数千种真实场景中的风格类别三元组。引入课程持续学习框架,在混合数据集(清洁与噪声)上训练,使模型能泛化至未见风格,同时保持精准内容保真。此外,设计视频到视频风格迁移模块,提升时序一致性与视觉质量。TeleStyle在风格相似性、内容一致性与美学质量三项核心评估指标上均达到领先水平。代码与预训练模型已开源。

原文摘要 · Abstract (English)

Content-preserving style transfer, generating stylized outputs based on content and style references, remains a significant challenge for Diffusion Transformers (DiTs) due to the inherent entanglement of content and style features in their internal representations. In this technical report, we present TeleStyle, a lightweight yet effective model for both image and video stylization. Built upon Qwen-Image-Edit, TeleStyle leverages the base model's robust capabilities in content preservation and style customization. To facilitate effective training, we curated a high-quality dataset of distinct specific styles and further synthesized triplets using thousands of diverse, in-the-wild style categories. We introduce a Curriculum Continual Learning framework to train TeleStyle on this hybrid dataset of clean (curated) and noisy (synthetic) triplets. This approach enables the model to generalize to unseen styles without compromising precise content fidelity. Additionally, we introduce a video-to-video stylization module to enhance temporal consistency and visual quality. TeleStyle achieves state-of-the-art performance across three core evaluation metrics: style similarity, content consistency, and aesthetic quality. Code and pre-trained models are available at https://github.com/Tele-AI/TeleStyle

风格迁移扩散模型视频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。