AirGS实现高质量4D高斯流实时传输,解决动态视频沉浸体验的带宽与质量难题。
AirGS: Real-Time 4D Gaussian Streaming for Free-Viewpoint Video Experiences
- 将4D高斯视频转为多通道2D流,智能识别关键帧提升重建质量
- 训练速度提升6倍,每帧传输量减少近50%,帧级PSNR稳定超30
- 适合需要低延迟、高保真自由视角视频的实时应用如直播与VR
自由视角视频(FVV)通过允许用户从任意角度观看场景,带来沉浸式体验。作为主流的FVV生成技术,4D高斯点阵(4DGS)使用随时间变化的3D高斯椭球建模动态场景,并通过快速光栅化实现高质量渲染。然而,现有4DGS方法在长序列中出现质量退化,且对带宽和存储开销要求高,限制了其在实时和大规模部署中的应用。为此,我们提出AirGS,一种面向流媒体优化的4DGS框架,重构训练与分发流程,实现高质量、低延迟的FVV体验。AirGS将高斯视频流转换为多通道2D格式,并智能识别关键帧以增强帧重建质量;结合时间一致性与膨胀损失,降低训练时间和表示规模。为支持高效通信,将4DGS传输建模为整数线性规划问题,设计轻量级剪枝等级选择算法,自适应地裁剪需传输的高斯更新,平衡重建质量与带宽消耗。大量实验表明,当场景变化时,AirGS使PSNR质量偏差降低超过20%,帧级PSNR始终维持在30以上,训练速度提升6倍,每帧传输量相比当前最优4DGS方案减少近50%。
原文摘要 · Abstract (English)
Free-viewpoint video (FVV) enables immersive viewing experiences by allowing users to view scenes from arbitrary perspectives. As a prominent reconstruction technique for FVV generation, 4D Gaussian Splatting (4DGS) models dynamic scenes with time-varying 3D Gaussian ellipsoids and achieves high-quality rendering via fast rasterization. However, existing 4DGS approaches suffer from quality degradation over long sequences and impose substantial bandwidth and storage overhead, limiting their applicability in real-time and wide-scale deployments. Therefore, we present AirGS, a streaming-optimized 4DGS framework that rearchitects the training and delivery pipeline to enable high-quality, low-latency FVV experiences. AirGS converts Gaussian video streams into multi-channel 2D formats and intelligently identifies keyframes to enhance frame reconstruction quality. It further combines temporal coherence with inflation loss to reduce training time and representation size. To support communication-efficient transmission, AirGS models 4DGS delivery as an integer linear programming problem and design a lightweight pruning level selection algorithm to adaptively prune the Gaussian updates to be transmitted, balancing reconstruction quality and bandwidth consumption. Extensive experiments demonstrate that AirGS reduces quality deviation in PSNR by more than 20% when scene changes, maintains frame-level PSNR consistently above 30, accelerates training by 6 times, reduces per-frame transmission size by nearly 50% compared to the SOTA 4DGS approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。