arXiv:2607.24403cs.CV2026-07

用生成式方法压缩3D高斯点云,提升低码率下视图一致性。

GenSplatCodec: Feed-Forward Gaussian Splatting Compression via One-Step Diffusion

论文配图:GenSplatCodec: Feed-Forward Gaussian Splatting Compression via One-Step Diffusion
图 1 · 摘自论文原文
  • 将压缩重构转化为几何引导的生成解码,统一结构与外观流。
  • 在低码率下实现更高质量、跨视角一致的新视图重建。
  • 适合需要高效传输3D场景的实时应用或移动设备使用。

前馈3D高斯点云(3DGS)可实现无需逐场景优化的可扩展场景重建,但产生的高斯点密集,存储与传输成本高。现有前馈压缩方法将解码视为确定性表示恢复,在低码率下因高频纹理和视图相关外观丢失而表现不佳。尽管生成模型具潜力,但作为独立后处理会割裂生成与传输的场景结构,影响跨视角一致性。为此,我们提出GenSplatCodec,一种统一的前馈高斯编码器,将低码率压缩重构为几何引导的生成解码。我们在双流框架中设计细节感知的前馈编码方案,将紧凑的高斯结构流与轻量级参考外观流结合。进一步提出几何引导的一步生成解码方法,通过分层几何控制联合利用解码后的结构与外观线索,重建高保真且视图一致的新视角。最后,我们开发三阶段优化策略,稳定统一编码器的学习,并使生成解码器适配编码器生成的结构与外观线索。多数据集上的大量实验表明,GenSplatCodec在率失真(RD)性能上持续优于现有方法。

原文摘要 · Abstract (English)

Feed-forward 3D Gaussian Splatting (3DGS) enables scalable scene reconstruction without per-scene optimization, yet produces dense Gaussians that are costly to store and transmit. Existing feed-forward Gaussian compression methods formulate decoding as deterministic representation recovery, which becomes inadequate at low bitrates when high-frequency textures and view-dependent appearance are discarded. Although generative models offer a promising alternative, using them as standalone post-processing decouples generation from the transmitted scene structure, thereby compromising cross-view consistency. To address these limitations, we propose GenSplatCodec, a unified feed-forward Gaussian codec that reformulates low-bitrate Gaussian compression as geometry-guided generative decoding. We present a detail-aware feed-forward Gaussian coding scheme within a dual-stream formulation, where the resulting compact Gaussian structural stream is complemented by a lightweight reference appearance stream. We further introduce a geometry-guided one-step generative decoding approach that jointly exploits decoded structural and appearance cues through hierarchical geometry control to reconstruct high-fidelity and view-consistent novel views. Finally, we develop a three-stage optimization strategy that stabilizes the learning of the unified codec and adapts the generative decoder to codec-derived structural and appearance cues. Extensive experiments across multiple datasets demonstrate that GenSplatCodec consistently achieves superior rate-distortion (RD) performance over existing methods.

3D重建高斯点云生成压缩视图一致

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。