让3D体积内容像2D视频一样高效压缩,支持自由视角播放。
CATRF: Codec-Adaptive TriPlane Radiance Fields for Volumetric Content Delivery

- 用标准编解码器(如JPEG/HEVC)闭环训练,让特征自动适应压缩失真。
- 在静态与动态数据集上均实现更优码率-质量平衡,压缩效率超越现有方法。
- 无需自研编码器,适合想快速部署低码率3D视频流的开发者。
体积媒体有望推动下一代内容分发应用,但带宽需求仍是关键瓶颈。隐式和混合体积表示虽可减小模型尺寸,但仍需精心编码才能达到类似2D视频的码率。我们提出CATRF,一种面向平面因子化辐射场的、标准编解码器闭环的压缩框架。训练时,将2D特征平面量化并打包为编解码友好画布,经标准编解码器(JPEG/VP9/HEVC/AV1)往返处理后,再解包并反量化特征以进行体渲染。采用直通估计器(STE)将非可微的标准编解码流程嵌入训练循环,使辐射场特征能直接适应真实客户端编解码器造成的失真,且不引入任何学习型编解码参数。在静态与动态基准测试中,CATRF始终优于无编解码感知及学习型编解码器闭环基线,在压缩效率和解码速度上也超越近期压缩版3DGS方法。结果表明,这为低码率、抗压缩的体积表示提供了实用路径,适用于自由视角视频流场景。
原文摘要 · Abstract (English)
Volumetric media promises next-generation content delivery applications, but its bandwidth demand remains a key bottleneck. Implicit and hybrid volumetric representations reduce model sizes, yet still require careful coding to reach 2D video-like bitrates. We present CATRF, a standard-codec-in-the-loop compression framework for plane-factorized radiance fields. During training, we quantize and pack 2D feature planes into codec-friendly canvases, run a standard codec roundtrip (JPEG/VP9/HEVC/AV1), then unpack and dequantize the decoded features before volume rendering. We use a straight-through estimator (STE) to insert the non-differentiable, standard codec pipeline into the training loop, allowing radiance-field features to adapt directly to the real, client-side codec distortions without introducing any learned codec parameters. On both static and dynamic benchmarks, CATRF consistently achieves a better rate-distortion trade-off over codec-agnostic and learned-codec-in-the-loop baselines, and also outperforms recent compressed 3DGS methods in both compression efficiency and decoding speed. These results highlight a practical path toward low-bitrate, compression-resilient volumetric representations for free-viewpoint video streaming.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。