用短视频训练,让模型学会长时序细节恢复,提升视频超分效果。
Small Clips, Big Gains: Learning Long-Range Refocused Temporal Information for Video Super-Resolution
- 在短片段上训练,通过时序传播学习长视频依赖关系。
- 引入注意力机制,选择性强化有用时序信息,提升帧间融合效果。
- 适配主流架构,性能领先且训练高效,适合长视频处理场景。
视频超分辨率(VSR)相比单图超分辨率能更好利用时序信息,尤其基于循环结构的VSR模型在推理时可捕捉长程时序依赖,实现更优细节重建。然而,如何有效学习长视频中的长期依赖仍是关键挑战。为此,我们提出LRTI-VSR——一种新型递归式VSR训练框架,高效利用长程重聚焦时序信息。该框架采用通用训练策略,在短视频片段上训练时引入长视频的时序传播特征;同时设计了融合内外帧的Transformer模块,通过注意力机制选择性关注有用时序信息,并在前馈网络中进一步优化帧间信息利用。我们在基于CNN与Transformer的VSR架构上评估了LRTI-VSR,通过大量消融实验验证各组件贡献。在长视频测试集上的实验表明,该方法达到当前最优性能,同时保持训练与计算效率。
原文摘要 · Abstract (English)
Video super-resolution (VSR) can achieve better performance compared to single image super-resolution by additionally leveraging temporal information. In particular, the recurrent-based VSR model exploits long-range temporal information during inference and achieves superior detail restoration. However, effectively learning these long-term dependencies within long videos remains a key challenge. To address this, we propose LRTI-VSR, a novel training framework for recurrent VSR that efficiently leverages Long-Range Refocused Temporal Information. Our framework includes a generic training strategy that utilizes temporal propagation features from long video clips while training on shorter video clips. Additionally, we introduce a refocused intra&inter-frame transformer block which allows the VSR model to selectively prioritize useful temporal information through its attention module while further improving inter-frame information utilization in the FFN module. We evaluate LRTI-VSR on both CNN and transformer-based VSR architectures, conducting extensive ablation studies to validate the contribution of each component. Experiments on long-video test sets demonstrate that LRTI-VSR achieves state-of-the-art performance while maintaining training and computational efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。