提出一种快速生成真实超分辨率图像的新方法,保留扩散模型的随机性。
Noise-Started One-Step Real-World Super-Resolution via LR-Conditioned SplitMeanFlow and GAN Refinement

- 基于条件分割均值流,实现从噪声到高清图的一步映射。
- 在单步推理下达到当前最佳感知质量,优于同类方法。
- 适合追求高效且逼真图像生成的研究者与应用开发者。
预训练的文本到图像(T2I)扩散模型因其噪声启动生成过程,在真实世界图像超分辨率(Real-ISR)中展现出强大潜力,能够实现逼真的纹理合成并捕捉超分辨率的多对一特性。然而,基于扩散的Real-ISR方法仍面临效率与质量之间的根本权衡。多步方法通过在低分辨率(LR)条件下逐步去噪随机高斯噪声生成高质量结果,但采样速度慢;近期的单步方法虽大幅提升效率,但通常以直接从低分辨率到高分辨率的重建替代噪声启动生成,削弱了随机性,限制了细节的真实感。为此,我们提出SMFSR,一种基于低分辨率条件分割均值流与GAN精炼的噪声启动单步Real-ISR框架。SMFSR保持扩散模型的随机噪声起点,并学习一个在低分辨率图像条件下直接从噪声到高分辨率图像的映射。为此,区间分割一致性将多步生成轨迹提炼为单一平均速度预测,实现高效单步生成。为弥补渐进式精炼机会的减少,我们进一步引入一个基于DINOv3的判别器和变分分数蒸馏的GAN精炼阶段,使生成结果在冻结的扩散教师指导下更贴近自然图像分布。大量实验表明,SMFSR在单步推理下实现了当前最优的感知质量,同时保持高速推理性能。
原文摘要 · Abstract (English)
Pre-trained text-to-image (T2I) diffusion models have shown strong potential for real-world image super-resolution (Real-ISR), owing to their noise-started generation process that enables realistic texture synthesis and captures the one-to-many nature of super-resolution. However, diffusion-based Real-ISR methods still face a fundamental efficiency-quality trade-off. Multi-step methods generate high-quality results by iteratively denoising random Gaussian noise under LR conditioning, but suffer from slow sampling. Recent one-step methods greatly improve efficiency, yet they typically replace noise-started generation with direct LR-to-HR restoration, which weakens stochasticity and limits realistic detail synthesis. To address this issue, we propose SMFSR, a noise-started one-step Real-ISR framework via LR-conditioned SplitMeanFlow and GAN refinement. SMFSR preserves the random-noise starting point of diffusion models and learns a direct noise-to-HR mapping conditioned on the LR image. To this end, Interval Splitting Consistency distills the multi-step generative trajectory into a single average-velocity prediction, enabling efficient one-step generation. To compensate for the reduced opportunity for progressive refinement, we further introduce a GAN refinement stage, where a DINOv3-based discriminator enhances realistic texture synthesis and variational score distillation aligns the generated outputs with the natural image distribution under a frozen diffusion teacher. Extensive experiments demonstrate that SMFSR achieves state-of-the-art perceptual quality among one-step diffusion-based Real-ISR methods while retaining fast single-step inference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。