针对视频压缩中高频细节丢失问题,提出频域感知神经表示方法
Frequency-aware Neural Representation for Videos
- 分离低频与高频成分,分阶段监督重建全局结构与纹理细节
- 动态注入高频信息,显著提升边缘和细粒度纹理还原效果
- 适合追求高保真视频压缩的科研与工业应用
隐式神经表示(INRs)在视频压缩领域展现出巨大潜力,但现有方法普遍存在频谱偏差问题,过度偏好低频成分,导致重建结果过平滑且率失真性能不佳。本文提出频率感知神经表示(FaNeRV),通过显式解耦低频与高频成分,实现高效且忠实的视频重建。FaNeRV引入多分辨率监督策略,分阶段引导网络逐步捕捉全局结构与精细纹理;为进一步增强高频重建能力,提出动态高频注入机制,自适应强调挑战性区域;同时设计频域分解网络模块,提升跨频段特征建模能力。在标准基准上的大量实验表明,FaNeRV显著优于现有先进INR方法,并在率失真性能上达到传统编码器的竞争力。
原文摘要 · Abstract (English)
Implicit Neural Representations (INRs) have emerged as a promising paradigm for video compression. However, existing INR-based frameworks typically suffer from inherent spectral bias, which favors low-frequency components and leads to over-smoothed reconstructions and suboptimal rate-distortion performance. In this paper, we propose FaNeRV, a Frequency-aware Neural Representation for videos, which explicitly decouples low- and high-frequency components to enable efficient and faithful video reconstruction. FaNeRV introduces a multi-resolution supervision strategy that guides the network to progressively capture global structures and fine-grained textures through staged supervision . To further enhance high-frequency reconstruction, we propose a dynamic high-frequency injection mechanism that adaptively emphasizes challenging regions. In addition, we design a frequency-decomposed network module to improve feature modeling across different spectral bands. Extensive experiments on standard benchmarks demonstrate that FaNeRV significantly outperforms state-of-the-art INR methods and achieves competitive rate-distortion performance against traditional codecs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。