提出PNVC框架,让神经视频压缩更实用,速度与质量双提升。
PNVC: Towards Practical INR-based Video Compression

- 结合自编码器与过拟合方案,设计新架构与位置嵌入机制。
- 低延迟模式下比HEVC节省35%以上码率,解码速度超20帧/秒。
- 适合追求高效率、低延迟的实时视频压缩应用开发。
神经视频压缩近年来在码率-质量性能上展现出与传统编码器竞争的潜力。然而,基于自编码器的方法存在解码复杂度高的问题,而基于隐式神经表示(INR)的模型则存在系统延迟大等缺陷,阻碍了其实际应用。本文针对实用化神经视频编码,提出新型INR编码框架PNVC,创新性地融合自编码器与过拟合解决方案。通过结构重参数化架构、分层质量控制、调制熵建模及尺度感知位置嵌入等多项设计优化,支持低延迟(LD)与随机访问(RA)配置。在低延迟模式下,相比HEVC HM 18.0实现近35%+的BD-rate节省,优于当前顶尖的INR编码器HiNeRV(多约10%),也超过VTM 20.0(多5%),同时保持1080p内容20+ FPS的解码速度。该工作为INR视频编码迈向实际部署迈出关键一步。源代码将公开供评估。
原文摘要 · Abstract (English)
Neural video compression has recently demonstrated significant potential to compete with conventional video codecs in terms of rate-quality performance. These learned video codecs are however associated with various issues related to decoding complexity (for autoencoder-based methods) and/or system delays (for implicit neural representation (INR) based models), which currently prevent them from being deployed in practical applications. In this paper, targeting a practical neural video codec, we propose a novel INR-based coding framework, PNVC, which innovatively combines autoencoder-based and overfitted solutions. Our approach benefits from several design innovations, including a new structural reparameterization-based architecture, hierarchical quality control, modulation-based entropy modeling, and scale-aware positional embedding. Supporting both low delay (LD) and random access (RA) configurations, PNVC outperforms existing INR-based codecs, achieving nearly 35%+ BD-rate savings against HEVC HM 18.0 (LD) - almost 10% more compared to one of the state-of-the-art INR-based codecs, HiNeRV and 5% more over VTM 20.0 (LD), while maintaining 20+ FPS decoding speeds for 1080p content. This represents an important step forward for INR-based video coding, moving it towards practical deployment. The source code will be available for public evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。