用扩散模型联合训练上下采样器,提升真实场景超分辨率效果。
Co-learning Single-Step Diffusion Upsampler and Downsampler with Two Discriminators and Distillation
- 通过双判别器和循环蒸馏,协同优化单步扩散上下采样器。
- 在真实世界与人脸超分任务中均达到顶尖保真度与视觉质量。
- 仅需一步推理,兼顾效率与真实退化模拟能力,适合实际应用。
超分辨率(SR)旨在从低分辨率(LR)图像重建高分辨率(HR)图像,通常依赖有效的下采样生成多样且真实的训练对。本文提出一种协同学习框架,联合优化基于扩散的单步上采样器与可学习下采样器,引入两个判别器和循环蒸馏策略。所提出的可学习下采样器能更好地捕捉真实退化模式,同时保留低分辨率域中的结构细节,对提升超分辨率性能至关重要。通过扩散方法在训练中生成多样化的低-高分辨率图像对,使模型在不同退化条件下具备鲁棒性。我们在通用真实世界及特定领域人脸超分任务上验证了该方法的有效性,实现了顶尖的保真度与感知质量。该方法不仅以单次推理实现高效,还确保高质量重建,弥合了合成与真实超分辨率场景之间的差距。
原文摘要 · Abstract (English)
Super-resolution (SR) aims to reconstruct high-resolution (HR) images from their low-resolution (LR) counterparts, often relying on effective downsampling to generate diverse and realistic training pairs. In this work, we propose a co-learning framework that jointly optimizes a single-step diffusion-based upsampler and a learnable downsampler, enhanced by two discriminators and a cyclic distillation strategy. Our learnable downsampler is designed to better capture realistic degradation patterns while preserving structural details in the LR domain, which is crucial for enhancing SR performance. By leveraging a diffusion-based approach, our model generates diverse LR-HR pairs during training, enabling robust learning across varying degradations. We demonstrate the effectiveness of our method on both general real-world and domain-specific face SR tasks, achieving state-of-the-art performance in both fidelity and perceptual quality. Our approach not only improves efficiency with a single inference step but also ensures high-quality image reconstruction, bridging the gap between synthetic and real-world SR scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。