arXiv:2503.20349cs.CV2025-03ICCV被引 17

无需预训练模型,一步生成逼真超分辨率图像。

Consistency Trajectory Matching for One-Step Generative Super-Resolution

  • 通过概率流微分方程构建从低清到高清的确定性映射
  • 直接学习单步映射,避免教师模型性能限制
  • 通过轨迹对齐提升结果真实感,适合追求高效生成的场景

当前基于扩散模型的超分辨率方法虽效果优异,但推理开销大。现有蒸馏方法虽能加速至单步,却显著增加训练成本并受限于教师模型性能。为此,本文提出无需蒸馏的超分辨率一致性轨迹匹配方法(CTMSR),可在单步内生成逼真超分辨率图像。我们首先建立概率流常微分方程(PF-ODE)轨迹,实现从含噪低分辨率(LR)图像到高分辨率(HR)图像的确定性映射;随后采用一致性训练(CT)策略,直接学习该单步映射,无需依赖预训练扩散模型。为进一步提升性能并更好利用真实图像分布,我们设计分布轨迹匹配(DTM)损失,最小化生成结果与自然图像在PF-ODE轨迹上的差异,从而增强重建图像的真实感。大量实验表明,该方法在合成与真实数据集上均达到可比或更优性能,同时保持极低推理延迟。

原文摘要 · Abstract (English)

Current diffusion-based super-resolution (SR) approaches achieve commendable performance at the cost of high inference overhead. Therefore, distillation techniques are utilized to accelerate the multi-step teacher model into one-step student model. Nevertheless, these methods significantly raise training costs and constrain the performance of the student model by the teacher model. To overcome these tough challenges, we propose Consistency Trajectory Matching for Super-Resolution (CTMSR), a distillation-free strategy that is able to generate photo-realistic SR results in one step. Concretely, we first formulate a Probability Flow Ordinary Differential Equation (PF-ODE) trajectory to establish a deterministic mapping from low-resolution (LR) images with noise to high-resolution (HR) images. Then we apply the Consistency Training (CT) strategy to directly learn the mapping in one step, eliminating the necessity of pre-trained diffusion model. To further enhance the performance and better leverage the ground-truth during the training process, we aim to align the distribution of SR results more closely with that of the natural images. To this end, we propose to minimize the discrepancy between their respective PF-ODE trajectories from the LR image distribution by our meticulously designed Distribution Trajectory Matching (DTM) loss, resulting in improved realism of our recovered HR images. Comprehensive experimental results demonstrate that the proposed methods can attain comparable or even superior capabilities on both synthetic and real datasets while maintaining minimal inference latency.

超分辨率扩散模型单步生成一致性训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。