用大模型生成140万张风格统一且多样的图像,提升风格迁移效果。
MegaStyle: Constructing Diverse and Scalable Style Dataset via Consistent Text-to-Image Style Mapping

- 利用文本到图像的风格一致性生成大量风格数据。
- 构建140万张图像的数据集,支持风格编码与迁移任务。
- 适合研究风格迁移、数据构建的学者和开发者使用。
本文提出MegaStyle,一种新颖且可扩展的数据构建流程,用于创建具有内部风格一致性、外部风格多样性及高质量的风格数据集。通过利用当前大型生成模型在给定风格描述下生成同一风格图像的能力,我们构建了一个包含17万条风格提示和40万条内容提示的多样化提示库,并通过内容-风格提示组合生成大规模风格数据集MegaStyle-1.4M。基于该数据集,我们提出风格监督对比学习方法,微调风格编码器MegaStyle-Encoder以提取表达性强的风格特异性表征,并训练了基于FLUX的风格迁移模型MegaStyle-FLUX。大量实验表明,保持风格内一致性、风格间多样性与高质量对风格数据集至关重要,所提出的MegaStyle-1.4M也展现出显著有效性。在该数据集上训练的MegaStyle-Encoder与MegaStyle-FLUX可提供可靠的风格相似性度量和泛化能力强的风格迁移,为风格迁移领域做出重要贡献。更多结果见项目网站:https://jeoyal.github.io/MegaStyle/
原文摘要 · Abstract (English)
In this paper, we introduce MegaStyle, a novel and scalable data curation pipeline that constructs an intra-style consistent, inter-style diverse and high-quality style dataset. We achieve this by leveraging the consistent text-to-image style mapping capability of current large generative models, which can generate images in the same style from a given style description. Building on this foundation, we curate a diverse and balanced prompt gallery with 170K style prompts and 400K content prompts, and generate a large-scale style dataset MegaStyle-1.4M via content-style prompt combinations. With MegaStyle-1.4M, we propose style-supervised contrastive learning to fine-tune a style encoder MegaStyle-Encoder for extracting expressive, style-specific representations, and we also train a FLUX-based style transfer model MegaStyle-FLUX. Extensive experiments demonstrate the importance of maintaining intra-style consistency, inter-style diversity and high-quality for style dataset, as well as the effectiveness of the proposed MegaStyle-1.4M. Moreover, when trained on MegaStyle-1.4M, MegaStyle-Encoder and MegaStyle-FLUX provide reliable style similarity measurement and generalizable style transfer, making a significant contribution to the style transfer community. More results are available at our project website https://jeoyal.github.io/MegaStyle/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。