通过可逆融合让风格与内容自动分离,无需显式标注。
SCFlow: Implicitly Learning Style and Content Disentanglement with Flow Models
- 用流匹配学习风格与内容的双向可逆映射。
- 在51万样本数据集上训练,实现零样本生成与图像迁移。
- 适合需要解耦生成的视觉任务,如艺术风格迁移。
视觉模型中显式分离风格与内容仍具挑战,因其语义重叠和人类感知主观性。现有方法依赖生成或判别目标进行分离,但仍受概念纠缠固有模糊性的限制。本文提出新思路:能否跳过显式分离,转而学习风格与内容的可逆融合,使分离自然涌现?我们提出SCFlow,一种基于流匹配的框架,学习纠缠与解耦表示间的双向映射。关键洞察包括:1)仅训练融合风格与内容(明确任务),即可实现无监督的可逆解耦;2)流匹配支持任意分布,避免扩散模型与归一化流对高斯先验的依赖;3)构建包含51种风格×10,000个内容样本的合成数据集(共51万样本),系统模拟风格-内容配对以促进解耦。除可控生成任务外,SCFlow在ImageNet-1k和WikiArt上实现零样本泛化,表现竞争力,表明解耦可通过可逆融合过程自然产生。
原文摘要 · Abstract (English)
Explicitly disentangling style and content in vision models remains challenging due to their semantic overlap and the subjectivity of human perception. Existing methods propose separation through generative or discriminative objectives, but they still face the inherent ambiguity of disentangling intertwined concepts. Instead, we ask: Can we bypass explicit disentanglement by learning to merge style and content invertibly, allowing separation to emerge naturally? We propose SCFlow, a flow-matching framework that learns bidirectional mappings between entangled and disentangled representations. Our approach is built upon three key insights: 1) Training solely to merge style and content, a well-defined task, enables invertible disentanglement without explicit supervision; 2) flow matching bridges on arbitrary distributions, avoiding the restrictive Gaussian priors of diffusion models and normalizing flows; and 3) a synthetic dataset of 510,000 samples (51 styles $\times$ 10,000 content samples) was curated to simulate disentanglement through systematic style-content pairing. Beyond controllable generation tasks, we demonstrate that SCFlow generalizes to ImageNet-1k and WikiArt in zero-shot settings and achieves competitive performance, highlighting that disentanglement naturally emerges from the invertible merging process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。