arXiv:2506.01037cs.CV2025-06CVPR被引 12

用自监督ControlNet和时空Mamba提升真实视频超分质量

Self-supervised ControlNet with Spatio-Temporal Mamba for Real-world Video Super-resolution

  • 引入自监督ControlNet,利用高分辨率特征引导低分辨率视频
  • 通过3D选择性扫描实现高效时空注意力,保持帧间一致性
  • 三阶段训练策略稳定模型,显著减少生成伪影

现有基于扩散模型的视频超分辨率方法因固有随机性,易在高分辨率视频中引入复杂退化和明显伪影。本文提出一种抗噪的实时视频超分辨率框架,将自监督学习与Mamba结构融入预训练潜空间扩散模型。为保证相邻帧内容一致性,引入基于3D选择性扫描模块的视频状态空间块,构建全局时空注意力机制,在计算成本可控的前提下增强时序连贯性。为减少生成细节中的伪影,设计自监督ControlNet,利用高分辨率特征作为指导,通过对比学习从低分辨率视频中提取对退化不敏感的特征。最后提出基于高/低分辨率视频混合的三阶段训练策略以稳定训练过程。所提方法在真实视频超分辨率基准数据集上优于现有最先进方法,验证了模型设计与训练策略的有效性。

原文摘要 · Abstract (English)

Existing diffusion-based video super-resolution (VSR) methods are susceptible to introducing complex degradations and noticeable artifacts into high-resolution videos due to their inherent randomness. In this paper, we propose a noise-robust real-world VSR framework by incorporating self-supervised learning and Mamba into pre-trained latent diffusion models. To ensure content consistency across adjacent frames, we enhance the diffusion model with a global spatio-temporal attention mechanism using the Video State-Space block with a 3D Selective Scan module, which reinforces coherence at an affordable computational cost. To further reduce artifacts in generated details, we introduce a self-supervised ControlNet that leverages HR features as guidance and employs contrastive learning to extract degradation-insensitive features from LR videos. Finally, a three-stage training strategy based on a mixture of HR-LR videos is proposed to stabilize VSR training. The proposed Self-supervised ControlNet with Spatio-Temporal Continuous Mamba based VSR algorithm achieves superior perceptual quality than state-of-the-arts on real-world VSR benchmark datasets, validating the effectiveness of the proposed model design and training strategies.

视频超分扩散模型自监督Mamba

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。