无需运动补偿的双轴时空注意力模型,提升真实场景视频超分辨率质量
DualX-VSR: Dual Axial Spatial$\times$Temporal Transformer for Real-World Video Super-Resolution without Motion Compensation
- 采用双轴时空注意力机制,分方向融合空间与时间信息
- 在Real-World VSR数据集上达到41.39dB PSNR,超越现有方法
- 无需光流对齐,适合复杂真实视频场景
基于Transformer的模型如ViViT和TimeSformer在建模时空依赖方面取得进展。近期视频生成模型如Sora和Vidu进一步展示了Transformer在长程特征提取与整体时空建模中的潜力。然而,直接将这些模型应用于真实世界视频超分辨率(VSR)面临挑战,因为VSR要求像素级精度,而分块与序列注意力机制可能破坏此精度。尽管最近的Transformer-VSR模型尝试通过小块和局部注意力缓解问题,但仍受限于感受野不足及依赖光学流对齐,后者在真实场景中易引入误差。为此,本文提出Dual Axial Spatial×Temporal Transformer(DualX-VSR),引入一种新颖的双轴时空注意力机制,沿正交方向整合空间与时间信息,无需运动补偿,结构简化且实现连贯的时空表征。结果表明,DualX-VSR在真实世界VSR任务中实现了高保真度与优异性能。
原文摘要 · Abstract (English)
Transformer-based models like ViViT and TimeSformer have advanced video understanding by effectively modeling spatiotemporal dependencies. Recent video generation models, such as Sora and Vidu, further highlight the power of transformers in long-range feature extraction and holistic spatiotemporal modeling. However, directly applying these models to real-world video super-resolution (VSR) is challenging, as VSR demands pixel-level precision, which can be compromised by tokenization and sequential attention mechanisms. While recent transformer-based VSR models attempt to address these issues using smaller patches and local attention, they still face limitations such as restricted receptive fields and dependence on optical flow-based alignment, which can introduce inaccuracies in real-world settings. To overcome these issues, we propose Dual Axial Spatial$\times$Temporal Transformer for Real-World Video Super-Resolution (DualX-VSR), which introduces a novel dual axial spatial$\times$temporal attention mechanism that integrates spatial and temporal information along orthogonal directions. DualX-VSR eliminates the need for motion compensation, offering a simplified structure that provides a cohesive representation of spatiotemporal information. As a result, DualX-VSR achieves high fidelity and superior performance in real-world VSR task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。