用扩散Transformer实现超高清艺术风格迁移,效果更自然无瑕疵。
U-StyDiT: Ultra-high Quality Artistic Style Transfer Using Diffusion Transformers
- 基于扩散Transformer构建,通过多视角调制学习风格信息
- 在自建的40万张图像数据集Aes4M上训练,生成图像质量显著提升
- 适合追求高画质艺术创作的设计师与研究人员
超高清艺术风格迁移旨在使用风格图像中学习到的风格信息重绘高质量内容图像。现有方法分为基于风格重建和内容-风格解耦两类,虽能生成部分艺术化图像,但仍存在明显伪影和不协调图案,难以生成超高清艺术风格图像。为此,本文提出U-StyDiT方法,基于基于Transformer的扩散模型(DiT),学习内容-风格解耦,生成超高清艺术风格图像。首先设计多视角风格调制器(MSM),从局部与全局角度学习风格信息,并引导模型生成符合该风格的图像;其次引入StyDiT模块,同步学习内容与风格条件。此外,构建包含10个类别、每类40万张图像的超高清艺术图像数据集Aes4M,解决了现有方法因数据集规模小、图像质量低导致生成效果不佳的问题。大量定性与定量实验表明,U-StyDiT在生成质量上优于当前最优艺术风格迁移方法。据我们所知,这是首个基于扩散Transformer实现超高清艺术风格迁移的方法。
原文摘要 · Abstract (English)
Ultra-high quality artistic style transfer refers to repainting an ultra-high quality content image using the style information learned from the style image. Existing artistic style transfer methods can be categorized into style reconstruction-based and content-style disentanglement-based style transfer approaches. Although these methods can generate some artistic stylized images, they still exhibit obvious artifacts and disharmonious patterns, which hinder their ability to produce ultra-high quality artistic stylized images. To address these issues, we propose a novel artistic image style transfer method, U-StyDiT, which is built on transformer-based diffusion (DiT) and learns content-style disentanglement, generating ultra-high quality artistic stylized images. Specifically, we first design a Multi-view Style Modulator (MSM) to learn style information from a style image from local and global perspectives, conditioning U-StyDiT to generate stylized images with the learned style information. Then, we introduce a StyDiT Block to learn content and style conditions simultaneously from a style image. Additionally, we propose an ultra-high quality artistic image dataset, Aes4M, comprising 10 categories, each containing 400,000 style images. This dataset effectively solves the problem that the existing style transfer methods cannot produce high-quality artistic stylized images due to the size of the dataset and the quality of the images in the dataset. Finally, the extensive qualitative and quantitative experiments validate that our U-StyDiT can create higher quality stylized images compared to state-of-the-art artistic style transfer methods. To our knowledge, our proposed method is the first to address the generation of ultra-high quality stylized images using transformer-based diffusion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。