端到端融合3D高斯喷溅与语义编码,实现实时沉浸式视频通信。
Generalizable 3D Gaussian Splatting enabled Semantic Coding for Real-Time Immersive Video Communications

- 将3D高斯喷溅重建与语义编码统一为端到端框架,避免冗余计算。
- 支持双视角并行处理,压缩效率提升且延迟更低。
- 适用于真实世界多场景,对压缩伪影鲁棒性强,适合实际部署。
实时沉浸式视频通信,尤其是高保真3D远程呈现,需要在动态场景实时重建与高效数据传输之间取得平衡。尽管前馈式3D高斯喷溅(3DGS)已实现实时渲染,但将多视图视频编码与3D重建解耦会带来压缩效率低下和计算复杂度高的问题。为此,我们提出GS-SCNet,首个将通用3DGS重建与专用深度语义编码管道无缝集成的端到端框架。其核心技术贡献包括:(i) 提出视差引导的并行语义编解码器,利用极线几何先验通过视差补偿与语义融合实现跨视角上下文交互,支持双视图实时并行处理,显著提升率失真性能;(ii) 设计轻量级高斯参数预测器,直接将解码后的语义潜在变量映射为3DGS属性,无需中间像素域重建。通过将编解码器与任务专用预测器耦合,仅需一次提取几何相关性,有效消除传统解耦范式中的冗余计算瓶颈。在合成与真实人类数据集上的大量实验表明,GS-SCNet在压缩效率、渲染质量与实时性能间实现了更优权衡。尤其在跨域真实数据上表现出强泛化能力与对压缩伪影的鲁棒性,显著优于传统解耦传输范式。
原文摘要 · Abstract (English)
Real-time immersive video communications, particularly high-fidelity 3D telepresence, necessitates a synergistic balance between instantaneous dynamic scene reconstruction and high-efficiency data transmission. While recent advancements in feed-forward 3D Gaussian Splatting (3DGS) have enabled real-time rendering, performing multi-view video coding and 3D reconstruction in a decoupled manner leads to suboptimal compression efficiency and high computational complexity. To address this, we propose GS-SCNet, the first unified end-to-end framework that seamlessly integrates generalizable 3DGS reconstruction with a dedicated deep Semantic Coding pipeline. Our architecture is underpinned by two core technical contributions: (i) we introduce a Disparity-Guided Parallel Semantic Codec that exploits epipolar geometric priors to facilitate cross-view contextual interaction via disparity compensation and semantic fusion, thereby enabling real-time parallel processing of stereo streams while significantly enhancing rate-distortion performance, and (ii) we develop a Lightweight Gaussian Parameter Predictor which directly projects decoded semantic latents into 3DGS attributes, obviating the need for intermediate pixel-domain reconstruction. By coupling the codec with the task-specific predictor, our framework extracts geometric correlations only once, effectively eliminating the redundant computational bottleneck inherent in conventional decoupled paradigms. Extensive evaluations on both synthetic and real-world human datasets demonstrate that GS-SCNet achieves a superior trade-off across compression efficiency, rendering quality, and real-time performance. Notably, our framework exhibits strong cross-domain generalization and robustness against compression artifacts when applied to out-of-domain real-world data, significantly outperforming conventional decoupled transmission paradigms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。