arXiv:2502.01816cs.CVcs.LG2025-02CVPR被引 2

轻量级视频超分模型仅用230万参数达顶尖效果

Low-Resource Video Super-Resolution using Memory, Wavelets, and Deformable Convolutions

  • 融合小波与可变形卷积,提升边缘信息利用
  • 单个记忆张量实现帧间时序建模,降低计算开销
  • 适合边缘设备实时视频增强,如流媒体显示

视频超分辨率(VSR)在资源受限的边缘设备上部署时,重建质量与计算开销之间仍存在巨大权衡。尽管近年基于Transformer的VSR模型在重建质量上达到新高度,但其计算成本过高。而近期推出的轻量级模型仍难以达到顶尖性能。本文提出一种新型轻量级、参数高效的VSR神经架构,仅用230万参数即达到当前最优重建精度。模型通过多项架构设计提升信息利用率:首先,在卷积层间嵌入2D小波分解,利用视觉数据中边缘的空间稀疏性先验;其次,使用单一记忆张量捕获帧间时序信息,避免传统记忆机制带来的高计算成本;最后,采用残差可变形卷积,隐式对齐帧间物体,增强帧间特征差异中的空间信息。该架构为边缘设备实现实时视频超分辨率(如流媒体显示)提供了可行路径。

原文摘要 · Abstract (English)

The tradeoff between reconstruction quality and compute required for video super-resolution (VSR) remains a formidable challenge in its adoption for deployment on resource-constrained edge devices. While transformer-based VSR models have set new benchmarks for reconstruction quality in recent years, these require substantial computational resources. On the other hand, lightweight models that have been introduced even recently struggle to deliver state-of-the-art reconstruction. We propose a novel lightweight and parameter-efficient neural architecture for VSR that achieves state-of-the-art reconstruction accuracy with just 2.3 million parameters. Our model enhances information utilization based on several architectural attributes. Firstly, it uses 2D wavelet decompositions strategically interlayered with learnable convolutional layers to utilize the inductive prior of spatial sparsity of edges in visual data. Secondly, it uses a single memory tensor to capture inter-frame temporal information while avoiding the computational cost of previous memory-based schemes. Thirdly, it uses residual deformable convolutions for implicit inter-frame object alignment that improve upon deformable convolutions by enhancing spatial information in inter-frame feature differences. Architectural insights from our model can pave the way for real-time VSR on the edge, such as display devices for streaming data.

视频超分轻量模型小波变换可变形卷积

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。