一个模型搞定图像生成、分割和分类,效果领先。
Symmetrical Flow Matching: Unified Image Generation, Segmentation, and Classification with Score-Based Generative Models
- 对称流匹配联合学习正反向变换,保证双向一致性。
- 25步推理即实现单步分割与分类,FID低至7.0。
- 支持像素级和图像级标签灵活条件,适合多任务场景。
流匹配已成为学习分布间连续变换的强大框架,实现了高保真生成建模。本文提出对称流匹配(SymmFlow),统一了语义分割、分类与图像生成于单一模型。通过对称学习目标,联合建模前向与反向变换,确保双向一致性,并保留足够熵以支持生成多样性。引入新训练目标,显式保留流中的语义信息,实现高效采样同时保持结构完整性,支持单步完成分割与分类而无需迭代优化。不同于以往严格的一一映射设定,SymmFlow可灵活处理像素级与图像级标签。在多个基准测试中,该方法在语义图像合成上达到顶尖性能:在CelebAMask-HQ上取得11.9的FID,在COCO-Stuff上为7.0,仅需25次推理步骤。同时在语义分割任务上表现优异,在分类任务中也展现良好潜力。
原文摘要 · Abstract (English)
Flow Matching has emerged as a powerful framework for learning continuous transformations between distributions, enabling high-fidelity generative modeling. This work introduces Symmetrical Flow Matching (SymmFlow), a new formulation that unifies semantic segmentation, classification, and image generation within a single model. Using a symmetric learning objective, SymmFlow models forward and reverse transformations jointly, ensuring bi-directional consistency, while preserving sufficient entropy for generative diversity. A new training objective is introduced to explicitly retain semantic information across flows, featuring efficient sampling while preserving semantic structure, allowing for one-step segmentation and classification without iterative refinement. Unlike previous approaches that impose strict one-to-one mapping between masks and images, SymmFlow generalizes to flexible conditioning, supporting both pixel-level and image-level class labels. Experimental results on various benchmarks demonstrate that SymmFlow achieves state-of-the-art performance on semantic image synthesis, obtaining FID scores of 11.9 on CelebAMask-HQ and 7.0 on COCO-Stuff with only 25 inference steps. Additionally, it delivers competitive results on semantic segmentation and shows promising capabilities in classification tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。