分离视频高低频成分,提升压缩效率与画质重建
Neural Video Representation for Redundancy Reduction and Consistency Preservation
- 将视频帧分解为高低频分量分别建模
- 在96%视频上优于现有方法HNeRV
- 适合需要高压缩比与高保真重建的场景
隐式神经表示(INR)可将多种信号嵌入神经网络,近年来因其处理多类信号的灵活性受到关注。在视频领域,INR通过将视频信号直接嵌入网络实现压缩。传统方法要么使用帧时间索引,要么使用单帧提取特征作为网络输入。后者表达能力更强,但特征中常含冗余信息,违背压缩初衷,且影响高频细节重建。为此,本文提出一种新视频表示方法,分离重构帧的高低频成分:低频部分由时序信息生成,高频部分基于高频特征生成。实验表明,该方法在96%的视频上优于现有HNeRV方法,显著提升压缩性能与重建质量。
原文摘要 · Abstract (English)
Implicit neural representation (INR) embed various signals into neural networks. They have gained attention in recent years because of their versatility in handling diverse signal types. In the context of video, INR achieves video compression by embedding video signals directly into networks and compressing them. Conventional methods either use an index that expresses the time of the frame or features extracted from individual frames as network inputs. The latter method provides greater expressive capability as the input is specific to each video. However, the features extracted from frames often contain redundancy, which contradicts the purpose of video compression. Additionally, such redundancies make it challenging to accurately reconstruct high-frequency components in the frames. To address these problems, we focus on separating the high-frequency and low-frequency components of the reconstructed frame. We propose a video representation method that generates both the high-frequency and low-frequency components of the frame, using features extracted from the high-frequency components and temporal information, respectively. Experimental results demonstrate that our method outperforms the existing HNeRV method, achieving superior results in 96 percent of the videos.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。