arXiv:2603.26134cs.CV2026-03

轻量级扩散模型实现高效且稳定的视频超分辨率

InstaVSR: Taming Diffusion for Efficient and Temporally Consistent Video Super-Resolution

  • 剪枝单步扩散架构,去除冗余计算组件
  • 30帧2K×2K视频处理耗时<1分钟,显存仅7GB
  • 适合需要实时部署与流畅时序的视频增强场景

视频超分辨率(VSR)旨在从低分辨率输入重建高分辨率帧。尽管基于扩散的方法显著提升了感知质量,但其在视频应用中仍面临两大挑战:强生成先验易引发时序不稳定性,多帧扩散流程通常计算成本过高,难以实际部署。为此,我们提出InstaVSR,一种轻量级扩散框架,兼顾效率与时序一致性。InstaVSR融合三项设计:(1) 剪枝的单步扩散主干,移除传统扩散VSR流水线中的多个高开销组件;(2) 基于光流引导的时序正则化递归训练,提升帧间稳定性;(3) 在潜在空间与像素空间的双空间对抗学习,以在主干简化后仍保持高质量感知效果。在NVIDIA RTX 4090上,InstaVSR可在1分钟内处理30帧2K×2K分辨率视频,显存占用仅7GB,相比现有扩散方法显著降低计算成本,同时保持优异感知质量与更平滑的时序过渡。

原文摘要 · Abstract (English)

Video super-resolution (VSR) seeks to reconstruct high-resolution frames from low-resolution inputs. While diffusion-based methods have substantially improved perceptual quality, extending them to video remains challenging for two reasons: strong generative priors can introduce temporal instability, and multi-frame diffusion pipelines are often too expensive for practical deployment. To address both challenges simultaneously, we propose InstaVSR, a lightweight diffusion framework for efficient video super-resolution. InstaVSR combines three ingredients: (1) a pruned one-step diffusion backbone that removes several costly components from conventional diffusion-based VSR pipelines, (2) recurrent training with flow-guided temporal regularization to improve frame-to-frame stability, and (3) dual-space adversarial learning in latent and pixel spaces to preserve perceptual quality after backbone simplification. On an NVIDIA RTX 4090, InstaVSR processes a 30-frame video at 2K$\times$2K resolution in under one minute with only 7 GB of memory usage, substantially reducing the computational cost compared to existing diffusion-based methods while maintaining favorable perceptual quality with significantly smoother temporal transitions.

视频超分辨率扩散模型轻量化时序稳定

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。