提出新方法实现3D高斯点云20倍压缩,且保持高质量重建。
Feed-Forward 3D Gaussian Splatting Compression with Long-Context Modeling
- 用莫顿编码构建千级上下文结构,增强长程依赖建模。
- 实现20倍压缩率,性能优于现有通用压缩算法。
- 适合需要高效存储与传输3D场景的工业应用者。
3D高斯泼溅(3DGS)作为革命性的3D表示方法,其庞大的数据量成为广泛采用的主要障碍。尽管前馈式3DGS压缩为昂贵的逐场景训练压缩器提供了实用替代方案,但现有方法受限于变换编码网络的有限感受野和熵模型上下文容量,难以建模长程空间依赖。本文提出一种新型前馈式3DGS压缩框架,通过大规模上下文结构(基于莫顿编码的数千个高斯点)与细粒度空间-通道自回归熵模型,充分挖掘扩展上下文信息。同时设计基于注意力的变换编码模型,聚合远距离邻近高斯点特征以提取有信息量的潜在先验。该方法在前馈推理下实现3DGS高达20倍的压缩比,并在通用编码器中达到领先性能。
原文摘要 · Abstract (English)
3D Gaussian Splatting (3DGS) has emerged as a revolutionary 3D representation. However, its substantial data size poses a major barrier to widespread adoption. While feed-forward 3DGS compression offers a practical alternative to costly per-scene per-train compressors, existing methods struggle to model long-range spatial dependencies, due to the limited receptive field of transform coding networks and the inadequate context capacity in entropy models. In this work, we propose a novel feed-forward 3DGS compression framework that effectively models long-range correlations to enable highly compact and generalizable 3D representations. Central to our approach is a large-scale context structure that comprises thousands of Gaussians based on Morton serialization. We then design a fine-grained space-channel auto-regressive entropy model to fully leverage this expansive context. Furthermore, we develop an attention-based transform coding model to extract informative latent priors by aggregating features from a wide range of neighboring Gaussians. Our method yields a $20\times$ compression ratio for 3DGS in a feed-forward inference and achieves state-of-the-art performance among generalizable codecs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。