让视频超分一步完成,且画质更优、训练更稳。
DUO-VSR: Dual-Stream Distillation for One-Step Video Super-Resolution
- 双流蒸馏框架,融合分布匹配与对抗监督
- 三阶段训练,实现高保真一步生成
- 适合追求高效高质视频超分的开发者
基于扩散模型的视频超分辨率近期取得了卓越的视觉保真度,但仍面临采样成本过高的问题。虽然分布匹配蒸馏(DMD)可加速扩散模型实现一步生成,但直接应用于视频超分辨率常导致训练不稳定及监督不足。为此,本文提出DUO-VSR,一种基于双流蒸馏策略的三阶段框架,统一分布匹配与对抗监督以实现一步视频超分辨率。首先采用渐进式引导蒸馏初始化,通过保持轨迹一致性稳定后续训练;其次,双流蒸馏联合优化DMD与真实-虚假评分特征生成对抗网络(RFS-GAN)流,后者利用真实与虚假评分模型的判别特征提供互补对抗监督;最后,偏好引导精修阶段进一步对齐学生模型与感知质量偏好。大量实验表明,相较于此前一步视频超分方法,DUO-VSR在视觉质量和效率上均表现更优。
原文摘要 · Abstract (English)
Diffusion-based video super-resolution (VSR) has recently achieved remarkable fidelity but still suffers from prohibitive sampling costs. While distribution matching distillation (DMD) can accelerate diffusion models toward one-step generation, directly applying it to VSR often results in training instability alongside degraded and insufficient supervision. To address these issues, we propose DUO-VSR, a three-stage framework built upon a Dual-Stream Distillation strategy that unifies distribution matching and adversarial supervision for one-step VSR. Firstly, a Progressive Guided Distillation Initialization is employed to stabilize subsequent training through trajectory-preserving distillation. Next, the Dual-Stream Distillation jointly optimizes the DMD and Real-Fake Score Feature GAN (RFS-GAN) streams, with the latter providing complementary adversarial supervision leveraging discriminative features from both real and fake score models. Finally, a Preference-Guided Refinement stage further aligns the student with perceptual quality preferences. Extensive experiments demonstrate that DUO-VSR achieves superior visual quality and efficiency over previous one-step VSR approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。