arXiv:2505.00046eess.IVcs.CV2025-05被引 2

用超分辨率网络提升神经视频表示的细节还原能力

SR-NeRV: Improving Embedding Efficiency of Neural Video Representation via Super-Resolution

  • 将通用超分网络引入神经视频表示,分离细节重建任务
  • 在模型尺寸相近下,显著提升高频率细节的重建质量
  • 适合需要高质量压缩视频的应用场景

隐式神经表示(INRs)因其在多个领域建模复杂信号的能力而受到广泛关注。近期基于INR的框架在神经视频压缩中展现出潜力,能将视频内容嵌入紧凑的神经网络。然而,在严格限制模型规模的情况下,这些方法常难以重建高频细节,这在实际压缩场景中尤为关键。为解决此问题,我们提出一种集成通用超分辨率(SR)网络的INR视频表示框架。该设计基于观察:高频成分在帧间通常具有较低的时间冗余性。通过将精细细节的重建任务交由预训练于自然图像的专用超分网络处理,所提方法提升了视觉保真度。实验表明,该方法在保持相近模型尺寸的前提下,优于传统INR基线,在重建质量上表现更优。

原文摘要 · Abstract (English)

Implicit Neural Representations (INRs) have garnered significant attention for their ability to model complex signals in various domains. Recently, INR-based frameworks have shown promise in neural video compression by embedding video content into compact neural networks. However, these methods often struggle to reconstruct high-frequency details under stringent constraints on model size, which are critical in practical compression scenarios. To address this limitation, we propose an INR-based video representation framework that integrates a general-purpose super-resolution (SR) network. This design is motivated by the observation that high-frequency components tend to exhibit low temporal redundancy across frames. By offloading the reconstruction of fine details to a dedicated SR network pre-trained on natural images, the proposed method improves visual fidelity. Experimental results demonstrate that the proposed method outperforms conventional INR-based baselines in reconstruction quality, while maintaining a comparable model size.

神经视频超分辨率压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。