NVRC++用轻量神经表示实现跨复杂度的高效视频压缩。
Enhanced Neural Video Representation Compression across Extreme Complexity and Quality Scales

- 采用多分辨率特征网格的轻量神经表示,统一架构支持不同复杂度。
- 四档复杂度(7k~360k MACs/pixel)覆盖宽比特率范围,实现实时解码。
- 优化框架减少冗余,熵模型高效压缩参数,适合部署于资源受限设备。
隐式神经表示(INRs)近年来成为视频压缩的有前途方法,兼具优异的率失真性能与快速解码能力。然而,现有神经视频编解码器难以平衡复杂度与可扩展性:轻量模型在不同比特率下性能下降,高性能模型则因复杂度随质量提升而扩展受限。这种缺乏统一架构、无法在广泛比特率范围内保持恒定复杂度的问题,严重限制了其实际部署。为此,我们提出NVRC++,一种基于INR的新型视频编解码器,采用轻量级INR搭配多个高分辨率特征网格,在任意复杂度层级实现高可扩展性。同时引入优化框架,高效拟合长视频序列的高分辨率网格,利用时空冗余而无显著计算或内存开销。此外,设计先进的熵模型以高效压缩高维网格参数。实验表明,NVRC++提供四个复杂度等级(7kMACs/pixel至360kMACs/pixel),每个层级均覆盖广泛的比特率与质量范围,并支持实时解码。相比当前最优的INR视频编解码器NVRC,解码速度提升最高达7.6倍,且性能相当。
原文摘要 · Abstract (English)
Implicit neural representations (INRs) have recently emerged as a promising approach to video compression, delivering competitive rate-distortion performance alongside rapid decoding. However, existing neural video codecs struggle to balance complexity and scalability. Lightweight models often suffer from degraded compression performance when scaled to different bitrate/quality levels, whereas high-performance models exhibit limited scalability, as their model complexity typically increases with quality. This lack of a unified architecture capable of maintaining consistent complexity across a wide range of bitrates severely limits their diverse real-world deployment. To address these challenges, we introduce NVRC++, a novel INR-based video codec that utilizes a lightweight INR with multiple high-resolution feature grids, providing high scalability at any given complexity level. This is paired with an optimization framework that enables efficient overfitting on high-resolution grids for long video sequences, thereby exploiting spatio-temporal redundancies without prohibitive computational or memory overhead. Additionally, an advanced entropy model is designed for efficiently compressing the high-dimensional grid parameters. As a result, NVRC++ provides four complexity levels (from 7kMACs/pixel to 360kMACs/pixel), each spanning wide bitrate and quality ranges while supporting real-time decoding. The experimental results show that NVRC++ offers a much faster decoding speed (up to 7.6x) compared to the SOTA INR-based video codec, NVRC, while delivering comparable performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。