arXiv:2601.22158cs.CV2026-01被引 48

提出像素级均值流模型,实现无需潜空间的一步图像生成。

One-step Latent-free Image Generation with Pixel Mean Flows

  • 分离网络输出与损失空间设计,直接预测图像像素
  • 256x256下FID达2.22,512x512下为2.48,性能领先
  • 适合追求高效生成的视觉应用开发者

当前主流的扩散/流模型通常具备两个特点:多步采样和在潜空间中操作。近期进展已在各自方向取得突破,推动了一步式无潜空间生成的发展。本文进一步推进该目标,提出像素均值流(pMF)。核心思想是将网络输出空间与损失空间分开设计:网络目标设定在预设的低维图像流形上(即图像预测),损失则通过速度空间中的均值流定义。我们引入图像流形与平均速度场之间的简单变换。实验表明,pMF在ImageNet 256x256分辨率下达到2.22 FID,512x512分辨率下为2.48 FID,填补了该领域关键空白。本研究有望推动扩散/流模型生成能力边界进一步拓展。

原文摘要 · Abstract (English)

Modern diffusion/flow-based models for image generation typically exhibit two core characteristics: (i) using multi-step sampling, and (ii) operating in a latent space. Recent advances have made encouraging progress on each aspect individually, paving the way toward one-step diffusion/flow without latents. In this work, we take a further step towards this goal and propose "pixel MeanFlow" (pMF). Our core guideline is to formulate the network output space and the loss space separately. The network target is designed to be on a presumed low-dimensional image manifold (i.e., x-prediction), while the loss is defined via MeanFlow in the velocity space. We introduce a simple transformation between the image manifold and the average velocity field. In experiments, pMF achieves strong results for one-step latent-free generation on ImageNet at 256x256 resolution (2.22 FID) and 512x512 resolution (2.48 FID), filling a key missing piece in this regime. We hope that our study will further advance the boundaries of diffusion/flow-based generative models.

图像生成扩散模型一步生成无潜空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。