arXiv:2509.16507cs.CV2025-09被引 2

一歩で高精細動画復元、効率と品質の両立を実現

OS-DiffVSR: Towards One-step Latent Diffusion Model for High-detailed Real-world Video Super-Resolution

  • 一歩の潜在拡散モデルでリアルな動画復元を実現
  • 複数フレーム融合で時間的整合性を維持しちらつき低減
  • 多数ステップ不要で既存手法を超える品質を達成

最近,潜在拡散模型在真实世界视频超分辨率(VSR)任务中表现出色,可通过多步扩散过程从退化的低分辨率输入重建高质量视频。与图像超分辨率(ISR)相比,VSR方法需处理视频中每一帧,对推理效率提出了挑战。然而,基于扩散的VSR方法在视频质量与推理效率之间始终存在权衡。本文提出一种针对真实世界视频超分辨率的一步扩散模型——OS-DiffVSR。具体而言,设计了一种新颖的相邻帧对抗训练范式,显著提升了合成视频的质量;此外,提出多帧融合机制,以保持帧间时间一致性并减少视频闪烁。在多个主流VSR基准上的大量实验表明,OS-DiffVSR甚至能在仅需一步采样的情况下,达到优于需数十步采样的现有扩散式VSR方法的性能。

原文摘要 · Abstract (English)

Recently, latent diffusion models has demonstrated promising performance in real-world video super-resolution (VSR) task, which can reconstruct high-quality videos from distorted low-resolution input through multiple diffusion steps. Compared to image super-resolution (ISR), VSR methods needs to process each frame in a video, which poses challenges to its inference efficiency. However, video quality and inference efficiency have always been a trade-off for the diffusion-based VSR methods. In this work, we propose One-Step Diffusion model for real-world Video Super-Resolution, namely OS-DiffVSR. Specifically, we devise a novel adjacent frame adversarial training paradigm, which can significantly improve the quality of synthetic videos. Besides, we devise a multi-frame fusion mechanism to maintain inter-frame temporal consistency and reduce the flicker in video. Extensive experiments on several popular VSR benchmarks demonstrate that OS-DiffVSR can even achieve better quality than existing diffusion-based VSR methods that require dozens of sampling steps.

视频超分扩散模型一歩推論時系列一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。