arXiv:2608.04818cs.CV2026-08

提出新型像素级去噪框架,实现无需潜变量的一步生成

Rethinking Pixel Mean Flows via Interval Denoiser

论文配图:Rethinking Pixel Mean Flows via Interval Denoiser
图 1 · 摘自论文原文
  • 基于流匹配微分方程构建精确的中间状态映射
  • 单步生成FID达4.55,双步达3.98,优于现有方法
  • 适合追求高效生成的图像合成研究者

现代扩散模型与基于流的方法正趋向少步、无潜变量生成,以规避多步采样计算开销和外部自编码器的重建瓶颈。本文提出区间去噪器(Interval Denoiser),一个理论严谨的无潜变量生成框架。该框架直接源于流匹配微分方程,建立中间轨迹状态的精确解析映射。与以往方法不同,我们的预测被证明在任意时间区间内位于低维流形上,使网络直接在像素上进行回归成为可能。此外,通过避免经验代数替换,本公式正确分离纯时间导数,防止梯度偏差,确保一阶优化精确性。通过分析该目标函数,我们引入残差裁剪与时间采样课程策略,实现有效长区间训练并提升少步性能。在ImageNet 256x256上从零训练,单步(1-NFE)FID为4.55,双步(2-NFE)FID为3.98,无需感知损失。

原文摘要 · Abstract (English)

Modern diffusion and flow-based models are increasingly moving toward few-step, latent-free generation to bypass the computational overhead of multi-step sampling and the reconstruction bottlenecks of external autoencoders. We propose the Interval Denoiser, a theoretically rigorous framework for latent-free generation. Derived directly from the flow matching ODE, it establishes an exact analytical mapping for intermediate trajectory states. Unlike prior formulations, our prediction is shown to reside on a low-dimensional manifold across any time interval, making the regression tractable for a network operating directly on pixels. Furthermore, by avoiding empirical algebraic substitutions, our formulation correctly isolates the pure time derivative to prevent biased gradient evaluations and ensure exact first-order optimization. By analyzing this objective, we equip our framework with residual clipping and a time-sampling curriculum, enabling effective long-interval training and improving few-step performance. Trained from scratch on ImageNet 256x256, our model achieves an FID of 4.55 in one step (1-NFE) and 3.98 in two steps (2-NFE) without perceptual losses.

生成模型扩散模型少步生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。