arXiv:2412.04469cs.CVcs.AI2024-12NeurIPS被引 31

QUEEN通过量化与稀疏化动态高斯,实现超低延迟的自由视角视频流传输。

QUEEN: QUantized Efficient ENcoding of Dynamic Gaussians for Streaming Free-viewpoint Videos

  • 直接学习相邻帧间高斯属性残差,无需结构约束,提升重建质量
  • 每帧仅0.7MB模型大小,训练<5秒,渲染达350FPS,动态场景表现优异
  • 基于视图梯度差分离静态/动态内容,加速训练并优化压缩效率

在线自由视角视频(FVV)流传输是极具挑战性但研究较少的问题,需在实时约束下实现体积表示的增量更新、快速训练与渲染,同时保持小内存开销以利于高效传输。若能实现,可推动3D视频会议、实时体视频广播等新应用。本文提出一种基于3D高斯喷溅(3D-GS)的量化高效编码框架QUEEN,直接在每个时间步学习连续帧间的高斯属性残差,且不对残差施加结构约束,从而实现高质量重建和强泛化能力。为高效存储残差,设计了包含可学习潜在解码器和门控模块的量化-稀疏框架:前者用于对非位置属性残差进行有效量化,后者用于稀疏化位置残差。引入高斯视图空间梯度差向量作为信号,分离场景中的静态与动态内容,指导稀疏学习并加速训练。在多个FVV基准上,QUEEN在所有指标上均优于现有最先进在线方法。尤其在高度动态场景中,其模型大小降至每帧仅0.7MB,训练时间不足5秒,渲染速度达350FPS。

原文摘要 · Abstract (English)

Online free-viewpoint video (FVV) streaming is a challenging problem, which is relatively under-explored. It requires incremental on-the-fly updates to a volumetric representation, fast training and rendering to satisfy real-time constraints and a small memory footprint for efficient transmission. If achieved, it can enhance user experience by enabling novel applications, e.g., 3D video conferencing and live volumetric video broadcast, among others. In this work, we propose a novel framework for QUantized and Efficient ENcoding (QUEEN) for streaming FVV using 3D Gaussian Splatting (3D-GS). QUEEN directly learns Gaussian attribute residuals between consecutive frames at each time-step without imposing any structural constraints on them, allowing for high quality reconstruction and generalizability. To efficiently store the residuals, we further propose a quantization-sparsity framework, which contains a learned latent-decoder for effectively quantizing attribute residuals other than Gaussian positions and a learned gating module to sparsify position residuals. We propose to use the Gaussian viewspace gradient difference vector as a signal to separate the static and dynamic content of the scene. It acts as a guide for effective sparsity learning and speeds up training. On diverse FVV benchmarks, QUEEN outperforms the state-of-the-art online FVV methods on all metrics. Notably, for several highly dynamic scenes, it reduces the model size to just 0.7 MB per frame while training in under 5 sec and rendering at 350 FPS. Project website is at https://research.nvidia.com/labs/amri/projects/queen

自由视角视频3D高斯量化编码实时渲染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。