arXiv:2502.07381cs.CV2025-02被引 3

针对压缩视频超分难题,提出时空一致的扩散模型方法。

Spatial Degradation-Aware and Temporal Consistent Diffusion Model for Compressed Video Super-Resolution

  • 引入退化感知模块与时空注意力机制,提升压缩视频重建质量。
  • 在多个基准数据集上显著优于现有方法,有效恢复纹理细节。
  • 适合需要高质量视频增强的应用场景,如流媒体修复。

由于存储和带宽限制,互联网传输的视频常呈现低分辨率和压缩伪影。尽管视频超分辨率(VSR)是有效的视频增强技术,但现有方法较少关注压缩视频。直接应用通用VSR方法难以提升含压缩伪影的实际视频,尤其在低码率下帧高度压缩时,量化信息丢失导致纹理重建困难。最近,扩散模型在低层视觉任务中表现出色。我们利用预训练扩散模型的先验知识,提出一种新的压缩视频超分辨率方法。为缓解空间失真并增强时序一致性,设计了空间退化感知与时间一致(SDATC)扩散模型。具体地,引入退化控制模块(DCM)调节扩散输入,降低低质量帧噪声对生成阶段的影响;随后,扩散模型通过微调的压缩感知提示模块(CAPM)和时空注意力模块(STAM)进行去噪生成。CAPM动态编码压缩相关信息至提示,使采样过程适应不同退化程度;STAM将空间注意力扩展至时空维度,有效捕捉时序相关性。此外,在每一步去噪中采用光流对齐,提升输出视频流畅度。大量实验表明,所提模块在基准数据集上均有效提升了压缩视频重建性能。

原文摘要 · Abstract (English)

Due to storage and bandwidth limitations, videos transmitted over the Internet often exhibit low quality, characterized by low-resolution and compression artifacts. Although video super-resolution (VSR) is an efficient video enhancing technique, existing VSR methods focus less on compressed videos. Consequently, directly applying general VSR approaches fails to improve practical videos with compression artifacts, especially when frames are highly compressed at a low bit rate. The inevitable quantization information loss complicates the reconstruction of texture details. Recently, diffusion models have shown superior performance in low-level visual tasks. Leveraging the high-realism generation capability of diffusion models, we propose a novel method that exploits the priors of pre-trained diffusion models for compressed VSR. To mitigate spatial distortions and refine temporal consistency, we introduce a Spatial Degradation-Aware and Temporal Consistent (SDATC) diffusion model. Specifically, we incorporate a distortion control module (DCM) to modulate diffusion model inputs, thereby minimizing the impact of noise from low-quality frames on the generation stage. Subsequently, the diffusion model performs a denoising process to generate details, guided by a fine-tuned compression-aware prompt module (CAPM) and a spatio-temporal attention module (STAM). CAPM dynamically encodes compression-related information into prompts, enabling the sampling process to adapt to different degradation levels. Meanwhile, STAM extends the spatial attention mechanism into the spatio-temporal dimension, effectively capturing temporal correlations. Additionally, we utilize optical flow-based alignment during each denoising step to enhance the smoothness of output videos. Extensive experimental results on benchmark datasets demonstrate the effectiveness of our proposed modules in restoring compressed videos.

视频超分扩散模型压缩修复

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。