用自监督ControlNet和时空Mamba提升真实视频超分质量
Self-supervised ControlNet with Spatio-Temporal Mamba for Real-world Video Super-resolution
- 引入自监督ControlNet,利用高分辨率特征引导低分辨率视频
- 通过3D选择性扫描实现高效时空注意力,保持帧间一致性
- 三阶段训练策略稳定模型,显著减少生成伪影
现有基于扩散模型的视频超分辨率方法因固有随机性,易在高分辨率视频中引入复杂退化和明显伪影。本文提出一种抗噪的实时视频超分辨率框架,将自监督学习与Mamba结构融入预训练潜空间扩散模型。为保证相邻帧内容一致性,引入基于3D选择性扫描模块的视频状态空间块,构建全局时空注意力机制,在计算成本可控的前提下增强时序连贯性。为减少生成细节中的伪影,设计自监督ControlNet,利用高分辨率特征作为指导,通过对比学习从低分辨率视频中提取对退化不敏感的特征。最后提出基于高/低分辨率视频混合的三阶段训练策略以稳定训练过程。所提方法在真实视频超分辨率基准数据集上优于现有最先进方法,验证了模型设计与训练策略的有效性。
原文摘要 · Abstract (English)
Existing diffusion-based video super-resolution (VSR) methods are susceptible to introducing complex degradations and noticeable artifacts into high-resolution videos due to their inherent randomness. In this paper, we propose a noise-robust real-world VSR framework by incorporating self-supervised learning and Mamba into pre-trained latent diffusion models. To ensure content consistency across adjacent frames, we enhance the diffusion model with a global spatio-temporal attention mechanism using the Video State-Space block with a 3D Selective Scan module, which reinforces coherence at an affordable computational cost. To further reduce artifacts in generated details, we introduce a self-supervised ControlNet that leverages HR features as guidance and employs contrastive learning to extract degradation-insensitive features from LR videos. Finally, a three-stage training strategy based on a mixture of HR-LR videos is proposed to stabilize VSR training. The proposed Self-supervised ControlNet with Spatio-Temporal Continuous Mamba based VSR algorithm achieves superior perceptual quality than state-of-the-arts on real-world VSR benchmark datasets, validating the effectiveness of the proposed model design and training strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。