探索神经视频编码中码率、失真与复杂度的权衡关系
On the Rate-Distortion-Complexity Trade-offs of Neural Video Coding
- 采用条件残差编码机制,融合时空信息提升编码效率
- 高分辨率条件特征导致计算量和内存占用显著增加
- 针对现有方法的瓶颈,提出更优的平衡方案,适合模型优化研究者
本文深入探讨现代神经视频编码中的码率-失真-复杂度权衡问题。近年来,研究重点转向挖掘神经视频编码的潜力,条件自编码器已成为高效编码的主流方法,其核心是利用空间和时间信息进行条件编码。然而,近期研究表明,条件编码可能面临信息瓶颈,表现甚至劣于传统残差编码。为解决此问题,当前方法引入大量高分辨率特征作为条件信号,导致乘加操作数、内存占用和模型规模显著上升。以DCVC为基准,本文研究新兴的条件残差编码及其变体如何在码率、失真和复杂度之间取得更好平衡。
原文摘要 · Abstract (English)
This paper aims to delve into the rate-distortion-complexity trade-offs of modern neural video coding. Recent years have witnessed much research effort being focused on exploring the full potential of neural video coding. Conditional autoencoders have emerged as the mainstream approach to efficient neural video coding. The central theme of conditional autoencoders is to leverage both spatial and temporal information for better conditional coding. However, a recent study indicates that conditional coding may suffer from information bottlenecks, potentially performing worse than traditional residual coding. To address this issue, recent conditional coding methods incorporate a large number of high-resolution features as the condition signal, leading to a considerable increase in the number of multiply-accumulate operations, memory footprint, and model size. Taking DCVC as the common code base, we investigate how the newly proposed conditional residual coding, an emerging new school of thought, and its variants may strike a better balance among rate, distortion, and complexity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。