用图文参考统一实现精准色彩风格迁移,效果超越现有方法。
MRStyle: A Unified Framework for Color Style Transfer with Multi-Modality Reference
- 构建图像与文本共享的统一风格空间,支持跨模态风格迁移。
- 生成无瑕疵的3D查表,高分辨率处理内存占用低,风格一致性强。
- 文本引导迁移效率高,适合开放场景下的个性化风格定制。
本文提出MRStyle,一种基于多模态参考(图像与文本)的统一色彩风格迁移框架。为实现两模态统一风格特征空间,我们设计了名为IRStyle的神经网络,通过交互双映射网络与联合监督学习管道,生成用于图像参考的风格化3D查找表,实现无视觉伪影、高分辨率低内存占用、强风格一致性,即使在显著颜色变化下亦能保持稳定。对于文本参考,我们将Stable Diffusion先验的文本特征与IRStyle风格特征对齐,实现文本引导的色彩风格迁移(TRStyle),训练与推理均高效,取得显著的开放集文本引导迁移效果。大量图像与文本实验表明,本方法在定性与定量评估中均优于当前最先进水平。
原文摘要 · Abstract (English)
In this paper, we introduce MRStyle, a comprehensive framework that enables color style transfer using multi-modality reference, including image and text. To achieve a unified style feature space for both modalities, we first develop a neural network called IRStyle, which generates stylized 3D lookup tables for image reference. This is accomplished by integrating an interaction dual-mapping network with a combined supervised learning pipeline, resulting in three key benefits: elimination of visual artifacts, efficient handling of high-resolution images with low memory usage, and maintenance of style consistency even in situations with significant color style variations. For text reference, we align the text feature of stable diffusion priors with the style feature of our IRStyle to perform text-guided color style transfer (TRStyle). Our TRStyle method is highly efficient in both training and inference, producing notable open-set text-guided transfer results. Extensive experiments in both image and text settings demonstrate that our proposed method outperforms the state-of-the-art in both qualitative and quantitative evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。