提出新方法让扩散模型在复杂退化视频中实现高保真超分辨率
DiffVSR: Revealing an Effective Recipe for Taming Robust Video Super-Resolution Against Complex Degradations
- 分阶段训练缓解模型学习负担,提升复杂退化下的表现
- 在严重退化视频上显著优于现有方法,峰值信噪比提升1.5dB
- 无需额外训练即可保持时间一致性,适合真实场景应用
扩散模型在图像修复中表现优异,但在视频超分辨率(VSR)中面临保真度与时序一致性难以兼顾的挑战。我们的评估发现:现有方法在严重退化视频上持续失效——而这正是扩散模型生成能力最需发挥之处。根源在于现有方法同时需建模复杂退化分布、内容表征与时序关系,但高质量训练数据有限。为此,我们提出DiffVSR,采用渐进式学习策略(PLS),通过分阶段训练系统性分解学习负担,显著提升复杂退化场景下的性能。框架还引入交错潜在过渡(ILT)技术,在不增加训练开销的前提下保持良好时序一致性。实验表明,该方法在竞争方法表现不佳的场景中尤为突出,尤其在严重退化视频上效果显著。研究揭示:优化学习策略比单纯堆砌架构更为关键,是实现鲁棒真实视频超分辨率的核心路径。
原文摘要 · Abstract (English)
Diffusion models have demonstrated exceptional capabilities in image restoration, yet their application to video super-resolution (VSR) faces significant challenges in balancing fidelity with temporal consistency. Our evaluation reveals a critical gap: existing approaches consistently fail on severely degraded videos--precisely where diffusion models' generative capabilities are most needed. We identify that existing diffusion-based VSR methods struggle primarily because they face an overwhelming learning burden: simultaneously modeling complex degradation distributions, content representations, and temporal relationships with limited high-quality training data. To address this fundamental challenge, we present DiffVSR, featuring a Progressive Learning Strategy (PLS) that systematically decomposes this learning burden through staged training, enabling superior performance on complex degradations. Our framework additionally incorporates an Interweaved Latent Transition (ILT) technique that maintains competitive temporal consistency without additional training overhead. Experiments demonstrate that our approach excels in scenarios where competing methods struggle, particularly on severely degraded videos. Our work reveals that addressing the learning strategy, rather than focusing solely on architectural complexity, is the critical path toward robust real-world video super-resolution with diffusion models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。