解决快速运动下4D高斯点云丢失问题,提升动态场景重建质量。
Implicit 4D Gaussian Splatting for Fast Motion with Large Inter-Frame Displacements

- 用时空位置隐式网络学习高斯属性,避免显式建模时间位移
- 在CMU体育数据集上,篮球场景比最强基线高1.83分PSNR
- 轻量级网络实现高效推理,适合高速运动视频重建
近期的4D高斯点云(4DGS)方法在快速运动且帧间位移较大的情况下表现不佳,因高斯属性在训练中学习不充分,导致快速移动物体常被丢失。本文提出时空位置隐式网络4DGS(SPIN-4DGS),通过显式采集的时空位置来学习高斯属性,而非建模时间位移,从而在大帧间位移下实现更精准的点云渲染。为避免对所有时空位置显式优化属性带来的内存开销,采用轻量级前馈网络,基于光栅化重建损失进行训练。该方法学习共享表示,有效捕捉时空一致性,在挑战性运动场景下保持稳定高质量渲染。大量实验表明,SPIN-4DGS在大位移条件下始终具备更高保真度,尤其在CMU Panoptic数据集的体育场景中,其性能显著优于现有方法。例如,在篮球场景中,相比最强基线模型D3DGS,PSNR提升1.83分。
原文摘要 · Abstract (English)
Recent 4D Gaussian Splatting (4DGS) methods often fail under fast motion with large inter-frame displacements, where Gaussian attributes are poorly learned during training, and fast-moving objects are often lost from the reconstruction. In this work, we introduce Spatiotemporal Position Implicit Network for 4DGS, coined SPIN-4DGS, which learns Gaussian attributes from explicitly collected spatiotemporal positions rather than modeling temporal displacements, thereby enabling more faithful splatting under fast motions with large inter-frame displacements. To avoid the heavy memory overhead of explicitly optimizing attributes across all spatiotemporal positions, we instead predict them with a lightweight feed-forward network trained under a rasterization-based reconstruction loss. Consequently, SPIN-4DGS learns shared representations across Gaussians, effectively capturing spatiotemporal consistency and enabling stable high-quality Gaussian splatting even under challenging motions. Across extensive experiments, SPIN-4DGS consistently achieves higher fidelity under large displacements, with clear improvements in PSNR and SSIM on challenging sports scenes from the CMU Panoptic dataset. For example, SPIN-4DGS notably outperforms the strongest baseline, D3DGS, by achieving +1.83 higher PSNR on the Basketball scene.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。