arXiv:2411.06685cs.CVcs.AI2024-11被引 5

提升视频压缩细节还原,通过高频特征增强模型表现

High-Frequency Enhanced Hybrid Neural Representation for Video Compression

  • 引入小波高频编码器提取视频高频信息
  • 在UVG和Bunny数据集上显著改善细节保留与压缩效果
  • 适合关注视频压缩细节质量的研究者或工程师

神经视频表示(NeRV)通过将视频内容编码为神经网络,简化了视频编解码流程并实现了快速解码,是视频压缩的有前景方案。然而,现有方法忽略了重建视频缺乏高频细节的问题。本文提出一种高频增强型混合神经表示网络,旨在利用高频信息提升网络对细微结构的合成能力。具体地,设计了包含小波频率分解器(WFD)模块的高频编码器,生成高频特征嵌入;进一步提出高频特征调制(HFM)模块,利用提取的高频嵌入增强解码器的拟合过程;最后结合优化的谐波解码器块与动态加权频率损失函数,有效减少高频信息丢失。在Bunny和UVG数据集上的实验表明,该方法在细节保持和压缩性能上均优于现有方法。

原文摘要 · Abstract (English)

Neural Representations for Videos (NeRV) have simplified the video codec process and achieved swift decoding speeds by encoding video content into a neural network, presenting a promising solution for video compression. However, existing work overlooks the crucial issue that videos reconstructed by these methods lack high-frequency details. To address this problem, this paper introduces a High-Frequency Enhanced Hybrid Neural Representation Network. Our method focuses on leveraging high-frequency information to improve the synthesis of fine details by the network. Specifically, we design a wavelet high-frequency encoder that incorporates Wavelet Frequency Decomposer (WFD) blocks to generate high-frequency feature embeddings. Next, we design the High-Frequency Feature Modulation (HFM) block, which leverages the extracted high-frequency embeddings to enhance the fitting process of the decoder. Finally, with the refined Harmonic decoder block and a Dynamic Weighted Frequency Loss, we further reduce the potential loss of high-frequency information. Experiments on the Bunny and UVG datasets demonstrate that our method outperforms other methods, showing notable improvements in detail preservation and compression performance.

视频压缩神经表示高频增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。