用3D非局部块提升视频超分辨率,让画面更清晰连贯。
Super-Resolution Generative Adversarial Networks based Video Enhancement
- 引入3D非局部块捕捉时空关联,突破单图模型局限。
- 实验显示时间连贯性增强,纹理更锐利,伪影减少。
- 提供大中小三种模型,兼顾效果与推理效率。
本文提出一种基于生成对抗网络的视频超分辨率方法,将传统单图超分辨率生成对抗网络(SRGAN)扩展至处理时空数据。针对视频需保持时间连续性的需求,设计了融合3D非局部块的新框架,以捕获空间与时间维度的长程依赖关系。构建基于块学习与先进退化技术的训练流程,模拟真实视频条件,使模型能同时学习局部细节与全局结构,提升泛化能力与稳定性。提出两种变体:一个更大模型以追求性能,一个轻量模型注重效率。实验表明,该方法在时间连贯性、纹理清晰度和视觉伪影控制方面显著优于传统单图方法。本工作为视频增强任务提供了实用的学习型解决方案,适用于流媒体、游戏及数字修复等场景。
原文摘要 · Abstract (English)
This study introduces an enhanced approach to video super-resolution by extending ordinary Single-Image Super-Resolution (SISR) Super-Resolution Generative Adversarial Network (SRGAN) structure to handle spatio-temporal data. While SRGAN has proven effective for single-image enhancement, its design does not account for the temporal continuity required in video processing. To address this, a modified framework that incorporates 3D Non-Local Blocks is proposed, which is enabling the model to capture relationships across both spatial and temporal dimensions. An experimental training pipeline is developed, based on patch-wise learning and advanced data degradation techniques, to simulate real-world video conditions and learn from both local and global structures and details. This helps the model generalize better and maintain stability across varying video content while maintaining the general structure besides the pixel-wise correctness. Two model variants-one larger and one more lightweight-are presented to explore the trade-offs between performance and efficiency. The results demonstrate improved temporal coherence, sharper textures, and fewer visual artifacts compared to traditional single-image methods. This work contributes to the development of practical, learning-based solutions for video enhancement tasks, with potential applications in streaming, gaming, and digital restoration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。