提出首个4K视频超分辨率基准,用新损失函数和控制网络提升细节恢复能力。
RealisVSR: Detail-enhanced Diffusion for Real-World 4K Video Super-Resolution
- 引入一致性控制网与高频修正损失,增强复杂运动与纹理重建。
- 仅需5%-25%训练数据量,在4K场景下显著优于现有方法。
- 构建首个公开4K VSR数据集,推动真实世界超分评估标准升级。
视频超分辨率(VSR)借助扩散模型有效缓解了传统GAN方法的过平滑问题。然而,当前仍存在三大挑战:1)基础模型对时序动态建模不一致;2)在复杂真实退化下高频率细节恢复有限;3)缺乏对细节增强与4K超分的充分评估,因现有方法主要依赖720P数据集且细节不足。为此,我们提出RealisVSR,一种高频率细节增强的视频扩散模型,包含三项核心创新:1)将一致性保留控制网(CPC)集成至Wan2.1视频扩散模型,以建模平滑复杂的运动并抑制伪影;2)提出基于小波分解与HOG特征约束的高频修正扩散损失(HR-Loss),用于纹理恢复;3)构建首个公开的4K VSR基准数据集RealisVideo-4K,包含1,000对高清视频-文本配对。利用Wan2.1先进的时空引导能力,本方法仅需现有方法5%-25%的训练数据量。在REDS、SPMCS、UDM10、YouTube-HQ、VideoLQ及RealisVideo-720P等多个基准上进行大量实验,验证其在超高清场景下的显著优势。
原文摘要 · Abstract (English)
Video Super-Resolution (VSR) has achieved significant progress through diffusion models, effectively addressing the over-smoothing issues inherent in GAN-based methods. Despite recent advances, three critical challenges persist in VSR community: 1) Inconsistent modeling of temporal dynamics in foundational models; 2) limited high-frequency detail recovery under complex real-world degradations; and 3) insufficient evaluation of detail enhancement and 4K super-resolution, as current methods primarily rely on 720P datasets with inadequate details. To address these challenges, we propose RealisVSR, a high-frequency detail-enhanced video diffusion model with three core innovations: 1) Consistency Preserved ControlNet (CPC) architecture integrated with the Wan2.1 video diffusion to model the smooth and complex motions and suppress artifacts; 2) High-Frequency Rectified Diffusion Loss (HR-Loss) combining wavelet decomposition and HOG feature constraints for texture restoration; 3) RealisVideo-4K, the first public 4K VSR benchmark containing 1,000 high-definition video-text pairs. Leveraging the advanced spatio-temporal guidance of Wan2.1, our method requires only 5-25% of the training data volume compared to existing approaches. Extensive experiments on VSR benchmarks (REDS, SPMCS, UDM10, YouTube-HQ, VideoLQ, RealisVideo-720P) demonstrate our superiority, particularly in ultra-high-resolution scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。