arXiv:2605.06421cs.CVcs.LG2026-05被引 2

分频处理图像生成,提升像素空间扩散模型的效率与质量

FREPix: Frequency-Heterogeneous Flow Matching for Pixel-Space Image Generation

论文配图:FREPix: Frequency-Heterogeneous Flow Matching for Pixel-Space Image Generation
图 1 · 摘自论文原文
  • 将图像生成分解为高低频分量,分别设计传输路径
  • 在256×256下达到1.91 FID,低NFE阶段表现更优
  • 适合追求高效、高质量像素级生成的研究者

像素空间扩散模型因避免了变分自编码器(VAEs)引入的表征瓶颈,重新成为有前景的替代方案。然而,现有方法大多将图像生成视为频率同质过程,忽视了低频与高频成分在角色和学习动态上的差异。为此,我们提出FREPix——一种面向像素空间生成的频率异构流匹配框架。FREPix显式地将生成过程分解为低频与高频成分,为其分配独立的传输路径,使用因子化网络分别预测,并通过频率感知目标进行训练。由此,从粗到细的生成成为明确的设计原则而非隐含行为。在ImageNet类到图像生成任务中,FREPix在像素空间生成模型中表现具有竞争力,在256×256下达到1.91 FID,512×512下为2.38 FID,尤其在训练初期和低NFE(非迭代步数)条件下表现突出。

原文摘要 · Abstract (English)

Pixel-space diffusion has re-emerged as a promising alternative to latent-space generation because it avoids the representation bottleneck introduced by VAEs. Yet most existing methods still treat image generation as a frequency-homogeneous process, overlooking the distinct roles and learning dynamics of low- and high-frequency components. To address this, we propose FREPix, a FREquency-heterogeneous flow matching framework for Pixel-space image generation. FREPix explicitly decomposes generation into low- and high-frequency components, assigns them separate transport paths, predicts them with a factorized network, and trains them with a frequency-aware objective. In this way, coarse-to-fine generation becomes an explicit design principle rather than an implicit behavior. On ImageNet class-to-image generation, FREPix achieves competitive results among pixel-space generation models, reaching 1.91 FID at $256\times256$ and 2.38 FID at $512\times512$, with particularly strong performance in the early stages of training and in the low-NFE regime.

图像生成扩散模型像素空间分频建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。