arXiv:2607.03524cs.CV2026-07被引 2

用感知空间优化流匹配,4-8步生成高质量图像。

Perceptual Flow Matching for Few-Step Generative Modeling

论文配图:Perceptual Flow Matching for Few-Step Generative Modeling
图 1 · 摘自论文原文
  • 在预训练感知空间中进行流匹配,而非传统潜在空间。
  • 仅需4-8步即可达到35-50步的生成质量,减少伪影。
  • 无需教师模型或辅助网络,适配主流生成流程。

我们提出感知流匹配(PFM),一种用于少步生成的简单而高效框架。不同于传统在变分自编码器(VAE)潜在空间中进行速度回归的方式,PFM利用预训练感知模型,在感知特征空间中监督流匹配。这一简单改进显著提升了流匹配模型的少步生成能力,将采样步数从35-50步降至4-8步,同时保持生成质量。与现有加速和蒸馏方法不同,PFM无需教师模型或辅助得分网络,可仅通过微小修改融入标准流匹配训练流程。在图像生成、视频生成和图像编辑任务上的大量实验表明,PFM始终产生高质量结果,且伪影少于现有蒸馏方法。我们进一步发现,感知监督使回归最小化从均值导向转向模式导向,使预测偏向流形上的模式,在粗粒度采样下仍保持准确。结果表明,只要在合适的表示空间中训练,标准流匹配可自然生成高质量少步生成器。该洞察或能激发未来对高效生成建模中表示感知目标的研究。

原文摘要 · Abstract (English)

We propose Perceptual Flow Matching (PFM), a simple yet effective framework for few-step generation in flow-matching models. Rather than performing velocity regression in the conventional VAE latent space, PFM supervises flow matching in a perceptual feature space using pretrained perceptual models. This simple change substantially improves the few-step generation capability of flow-matching models, reducing the number of sampling steps from 35-50 to 4-8 while preserving generation quality. Unlike existing acceleration and distillation approaches, PFM requires neither teacher models nor auxiliary score networks and can be integrated into standard flow-matching training pipelines with minimal modifications. Extensive experiments on image generation, video generation, and image editing tasks demonstrate that PFM consistently produces high-quality results while producing fewer artifacts than existing distillation-based methods. We further show that perceptual supervision shifts the regression minimizer from mean-seeking to mode-seeking, biasing predictions toward on-manifold modes that remain accurate under coarse few-step integration. Our results reveal that standard flow-matching training can naturally yield high-quality few-step generators when supervised in an appropriate representation space. We hope this insight inspires future research into representation-aware objectives for efficient generative modeling.

流匹配少步生成感知空间高效生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。