用噪声激励提升静态视频压缩效率,实现超低码率下无损还原。
Enhancing Neural Video Compression of Static Scenes with Positive-Incentive Noise
- 将短期变化视为正向激励噪声,优化神经压缩模型。
- 静态场景压缩率仅0.009%,比基线低20.5%的BD速率。
- 适合安防监控等需长期存储与低带宽传输的场景。
静态场景视频(如监控流、视频通话)占用了大量存储和网络流量。传统编码标准与神经视频压缩(NVC)方法因未能有效利用时间冗余,以及训练与测试数据分布差异大而效率低下。尽管近期生成式压缩提升了视觉质量,但引入了不可接受的幻觉细节。为此,我们提出正向激励相机(PIC)框架,将短期时间变化重新解释为正向激励噪声,用于微调NVC模型。通过分离瞬时变化与恒定背景,结构化先验信息被内化至压缩模型中。推理时,不变成分只需极少信令,显著降低传输数据量,同时保持像素级保真度。实验表明,PIC在极低压缩率0.009%下实现视觉无损重建,相较DCVC-FM基线节省20.5%的Bjøntegaard delta(BD)率。该方法以计算换带宽,适用于恶劣网络条件下的鲁棒传输及监控录像的经济长期保存。
原文摘要 · Abstract (English)
Static scene videos, such as surveillance feeds and videotelephony streams, constitute a dominant share of storage consumption and network traffic. However, both traditional standardized codecs and neural video compression (NVC) methods struggle to encode these videos efficiently due to inadequate usage of temporal redundancy and severe distribution gaps between training and test data, respectively. While recent generative compression methods improve perceptual quality, they introduce hallucinated details that are unacceptable in authenticity-critical applications. To overcome these limitations, we propose a positive-incentive camera (PIC) framework for static scene videos, where short-term temporal changes are reinterpreted as positive-incentive noise to facilitate NVC model finetuning. By disentangling transient variations from the persistent background, structured prior information is internalized in the compression model. During inference, the invariant component requires minimal signaling, thus reducing data transmission while maintaining pixel-level fidelity. Experiment results show that PIC achieves visually lossless reconstruction for static scenes at an extremely low compression rate of 0.009%, while the DCVC-FM baseline requires 20.5% higher Bjøntegaard delta (BD) rate. Our method provides an effective solution to trade computation for bandwidth, enabling robust video transmission under adverse network conditions and economic long-term retention of surveillance footage.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。