提出率感知压缩框架,训练时直接优化体积视频的紧凑表示。
Rate-aware Compression for NeRF-based Volumetric Video
- 训练阶段引入隐式熵模型估算码率,实现精准压缩控制。
- 自适应量化使NeRF表示在低码率下仍保持高质量,人眼感知差异小。
- 适用于多种体积视频表示,显著优于现有方法,适合实时传输场景。
神经辐射场(NeRF)推动了3D体积视频技术发展,但其庞大的数据量给存储与传输带来挑战。现有方法通常在训练后压缩表示,导致训练与压缩分离。本文提出率感知压缩框架,在训练阶段直接学习紧凑的NeRF表示。针对体积视频,采用简单有效策略减少时间冗余;训练中使用隐式熵模型估计码率,并将该模型编码进比特流辅助解码,实现精确码率控制。进一步提出自适应量化策略,学习最优量化步长,通过率失真权衡优化表示。实验表明,该框架适用于多种表示,在HumanRF和ReRF数据集上显著降低存储量,失真微小,达到当前最佳率失真性能:相比TeTriRF方法,在HumanRF上实现约-80% BD-rate,在ReRF上实现约-60% BD-rate。
原文摘要 · Abstract (English)
The neural radiance fields (NeRF) have advanced the development of 3D volumetric video technology, but the large data volumes they involve pose significant challenges for storage and transmission. To address these problems, the existing solutions typically compress these NeRF representations after the training stage, leading to a separation between representation training and compression. In this paper, we try to directly learn a compact NeRF representation for volumetric video in the training stage based on the proposed rate-aware compression framework. Specifically, for volumetric video, we use a simple yet effective modeling strategy to reduce temporal redundancy for the NeRF representation. Then, during the training phase, an implicit entropy model is utilized to estimate the bitrate of the NeRF representation. This entropy model is then encoded into the bitstream to assist in the decoding of the NeRF representation. This approach enables precise bitrate estimation, thereby leading to a compact NeRF representation. Furthermore, we propose an adaptive quantization strategy and learn the optimal quantization step for the NeRF representations. Finally, the NeRF representation can be optimized by using the rate-distortion trade-off. Our proposed compression framework can be used for different representations and experimental results demonstrate that our approach significantly reduces the storage size with marginal distortion and achieves state-of-the-art rate-distortion performance for volumetric video on the HumanRF and ReRF datasets. Compared to the previous state-of-the-art method TeTriRF, we achieved an approximately -80% BD-rate on the HumanRF dataset and -60% BD-rate on the ReRF dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。