用单一模型实现可变码率的神经体积视频压缩,提升画质与效率。
VRVVC: Variable-Rate NeRF-Based Volumetric Video Compression
- 采用紧凑的三平面残差表示建模动态场景时间冗余。
- 通过可学习量化与小规模MLP熵模型支持多码率输出。
- 端到端训练+多率失真损失,单模型覆盖广泛码率范围。
基于神经辐射场(NeRF)的体积视频已革新视觉媒体,提供高度沉浸与交互的自由视角视频体验。然而其庞大的数据量对存储与传输带来挑战。现有方法通常独立优化表示与压缩,或仅针对单一固定率-失真(RD)权衡。本文提出VRVVC,一种新型端到端联合优化的可变码率体积视频压缩框架,仅用一个模型即可实现可变比特率,并保持优异的RD性能。具体而言,VRVVC引入紧凑的三平面隐式残差表示用于长时动态场景的帧间建模,有效降低时间冗余。进一步提出基于可学习量化与小型MLP熵模型的可变码率残差压缩方案,通过预设拉格朗日乘子控制所有潜在表示的量化误差。最后,结合渐进式训练策略与多率失真损失函数,优化整个框架。大量实验表明,VRVVC在单一模型内实现广泛可变码率,且在多个数据集上均优于现有方法的RD性能。
原文摘要 · Abstract (English)
Neural Radiance Field (NeRF)-based volumetric video has revolutionized visual media by delivering photorealistic Free-Viewpoint Video (FVV) experiences that provide audiences with unprecedented immersion and interactivity. However, the substantial data volumes pose significant challenges for storage and transmission. Existing solutions typically optimize NeRF representation and compression independently or focus on a single fixed rate-distortion (RD) tradeoff. In this paper, we propose VRVVC, a novel end-to-end joint optimization variable-rate framework for volumetric video compression that achieves variable bitrates using a single model while maintaining superior RD performance. Specifically, VRVVC introduces a compact tri-plane implicit residual representation for inter-frame modeling of long-duration dynamic scenes, effectively reducing temporal redundancy. We further propose a variable-rate residual representation compression scheme that leverages a learnable quantization and a tiny MLP-based entropy model. This approach enables variable bitrates through the utilization of predefined Lagrange multipliers to manage the quantization error of all latent representations. Finally, we present an end-to-end progressive training strategy combined with a multi-rate-distortion loss function to optimize the entire framework. Extensive experiments demonstrate that VRVVC achieves a wide range of variable bitrates within a single model and surpasses the RD performance of existing methods across various datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。