用对比学习提升生成模型的内容风格解耦能力
Contrastive-Augmented Flow Matching for Style-Content Disentanglement

- 在流匹配中引入对比正则化,让内容与风格分离
- 在多个数据集上显著提升内容/风格检索准确率
- 适合需要可控生成和分布外鲁棒性的研究者
学习将内容与风格分离的表示对可控生成和组合泛化至关重要。然而,仅以生成目标训练的扩散模型和基于流的模型常产生纠缠或错位的因子。为此,我们提出对比增强流匹配(CAtFM),将对比正则化整合到可逆流匹配框架中,以促进结构化的内容-风格表示。不约束中间潜在变量或速度场,而是在训练时对预测终点施加对比监督,强制传输分布间的语义一致性,同时允许解耦自然涌现,无需假设内容与风格为严格纯或完全因子化。主要实验在CLIP嵌入空间进行,并用冻结的DINO和ALIGN编码器验证。在合成数据、域内风格及真实世界基准(ImageNet、WikiArt、DomainNet、DTD)上,CAtFM提升了内容与风格检索性能,增强了嵌入聚类分离度,并在分布外任务上表现更强鲁棒性。整体而言,CAtFM提供了一种简单方法,将判别式约束与确定性传输结合,提升分布偏移下的解耦性和鲁棒性。
原文摘要 · Abstract (English)
Learning representations that separate content and style is crucial for controllable generation and compositional generalization. However, diffusion and flow-based models trained primarily with generative objectives often produce entangled or misaligned factors. To address this gap, we introduce Contrastive Augmented Flow Matching (CAtFM), a framework that integrates contrastive regularization into an invertible flow matching formulation to promote structured content-style representations. Rather than constraining intermediate latents or velocity fields, we apply contrastive supervision to predicted endpoints during training, enforcing semantic consistency across transported distributions while allowing disentanglement to emerge implicitly, without assuming strictly pure or fully factorized content and style representations. Our main experiments operate in the CLIP embedding space, with additional validation using frozen DINO and ALIGN encoders. Across synthetic data, in-domain styles, and real-world benchmarks (ImageNet, WikiArt, DomainNet, and DTD), CAtFM improves content and style retrieval, enhances embedding cluster separation, and achieves stronger open-set robustness compared to generative and discriminative baselines. Overall, CAtFM provides a simple way to couple discriminative constraints with deterministic transport, improving disentanglement and robustness under distribution shift.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。