通过反向去风格化生成训练数据,提升风格迁移质量与规模。
OmniStyle2: Learning to Stylize by Learning to Destylize
- 先去风格再还原,用去风格化构建大规模真实配对数据。
- 在DeStyle-350K数据集上训练,风格迁移效果优于现有方法。
- 适合需要高质量、大规模风格迁移训练数据的研究者。
本文提出一种可扩展的监督风格迁移范式:不直接学习如何风格化,而是学习去风格化,将艺术图像中的风格元素消除,恢复其自然内容,从而在大规模上生成真实且像素对齐的训练样本。为此,我们设计了多阶段渐进式去风格化框架DeStylePipe,从全局通用去风格化开始,逐步过渡到类别特定指令适配,最终针对复杂风格采用专用模型微调。整个流程中嵌入了基于思维链(Chain-of-Thought)推理的DestyleCoT-Filter,用于评估每一步的内容保留与风格去除效果,筛选出高质量样本并剔除低质配对。基于此框架,我们构建了包含35万对样本的大型数据集DeStyle-350K,涵盖多样艺术风格与其底层内容。同时提出BCS-Bench基准,兼顾内容通用性与风格多样性,支持系统评估。大量实验表明,基于DeStyle-350K训练的模型在风格迁移质量上表现更优,验证了去风格化作为可靠、可扩展的监督范式的有效性。
原文摘要 · Abstract (English)
This paper introduces a scalable paradigm for supervised style transfer by inverting the problem: instead of learning to stylize directly, we learn to destylize, reducing stylistic elements from artistic images to recover their natural counterparts and thereby producing authentic, pixel-aligned training pairs at scale. To realize this paradigm, we propose DeStylePipe, a progressive, multi-stage destylization framework that begins with global general destylization, advances to category-wise instruction adaptation, and ultimately deploys specialized model adaptation for complex styles that prompt engineering alone cannot handle. Tightly integrated into this pipeline, DestyleCoT-Filter employs Chain-of-Thought reasoning to assess content preservation and style removal at each stage, routing challenging samples forward while discarding persistently low-quality pairs. Built on this framework, we construct DeStyle-350K, a large-scale dataset aligning diverse artistic styles with their underlying content. We further introduce BCS-Bench, a benchmark featuring balanced content generality and style diversity for systematic evaluation. Extensive experiments demonstrate that models trained on DeStyle-350K achieve superior stylization quality, validating destylization as a reliable and scalable supervision paradigm for style transfer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。