用注意力传播提升视频超分辨率的时空一致性与细节真实感
DC-VSR: Spatially and Temporally Consistent Video Super-Resolution with Video Diffusion Prior

- 引入时空注意力传播机制,跨块传递信息保持一致性
- 在REDs、Udacity等数据集上优于现有方法,峰值达43.21dB
- 适合需要高质量视频修复与增强的应用场景
视频超分辨率(VSR)旨在从低分辨率(LR)视频中重建高分辨率(HR)视频。成功的VSR需恢复真实细节并保证时空一致性。近期基于扩散模型的方法虽能生成逼真纹理,但其固有的随机性与分块处理方式常导致时空不一致。本文提出DC-VSR,通过新颖的空间注意力传播(SAP)与时间注意力传播(TAP)机制,利用自注意力在时空块间传递信息,实现一致性重建。同时引入细节抑制自注意力引导(DSSAG)以增强高频细节。大量实验表明,DC-VSR在REDs、Udacity等数据集上均超越现有方法,取得更高质量且一致的视频超分辨率结果。
原文摘要 · Abstract (English)
Video super-resolution (VSR) aims to reconstruct a high-resolution (HR) video from a low-resolution (LR) counterpart. Achieving successful VSR requires producing realistic HR details and ensuring both spatial and temporal consistency. To restore realistic details, diffusion-based VSR approaches have recently been proposed. However, the inherent randomness of diffusion, combined with their tile-based approach, often leads to spatio-temporal inconsistencies. In this paper, we propose DC-VSR, a novel VSR approach to produce spatially and temporally consistent VSR results with realistic textures. To achieve spatial and temporal consistency, DC-VSR adopts a novel Spatial Attention Propagation (SAP) scheme and a Temporal Attention Propagation (TAP) scheme that propagate information across spatio-temporal tiles based on the self-attention mechanism. To enhance high-frequency details, we also introduce Detail-Suppression Self-Attention Guidance (DSSAG), a novel diffusion guidance scheme. Comprehensive experiments demonstrate that DC-VSR achieves spatially and temporally consistent, high-quality VSR results, outperforming previous approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。