提出一种高效视频超分模型,加速推理且无需额外信息。
FCA2: Frame Compression-Aware Autoencoder for Modular and Fast Compressed Video Super-Resolution
- 借鉴高光谱图像结构,设计压缩感知维度压缩策略
- 推理速度显著提升,性能达当前最优水平
- 模块化设计适配多种框架,适合实际部署
当前先进的压缩视频超分辨率(CVSR)模型存在推理时间长、训练流程复杂及依赖辅助信息等问题。随着帧率提高,帧间差异减小,传统逐帧信息利用方法难以满足现代视频超分辨率需求。为此,我们受高光谱图像(HSI)与视频数据在结构和统计特性上的相似性启发,提出一种基于压缩感知的维度压缩策略,有效降低计算复杂度,加快推理速度,并增强跨帧时序信息提取能力。所提模块化架构可无缝集成至现有视频超分辨率框架中,具备强适应性和可迁移性。实验表明,该方法性能达到或超越当前最先进水平,同时显著缩短推理时间。本工作为解决CVSR中的关键瓶颈提供了高效实用的技术路径。代码将公开于 https://github.com/handsomewzy/FCA2。
原文摘要 · Abstract (English)
State-of-the-art (SOTA) compressed video super-resolution (CVSR) models face persistent challenges, including prolonged inference time, complex training pipelines, and reliance on auxiliary information. As video frame rates continue to increase, the diminishing inter-frame differences further expose the limitations of traditional frame-to-frame information exploitation methods, which are inadequate for addressing current video super-resolution (VSR) demands. To overcome these challenges, we propose an efficient and scalable solution inspired by the structural and statistical similarities between hyperspectral images (HSI) and video data. Our approach introduces a compression-driven dimensionality reduction strategy that reduces computational complexity, accelerates inference, and enhances the extraction of temporal information across frames. The proposed modular architecture is designed for seamless integration with existing VSR frameworks, ensuring strong adaptability and transferability across diverse applications. Experimental results demonstrate that our method achieves performance on par with, or surpassing, the current SOTA models, while significantly reducing inference time. By addressing key bottlenecks in CVSR, our work offers a practical and efficient pathway for advancing VSR technology. Our code will be publicly available at https://github.com/handsomewzy/FCA2.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。