用分层高斯点云实现高效视频建模,兼顾速度与一致性
GaussianVideo: Efficient Video Representation via Hierarchical Gaussian Splatting
- 分层学习策略逐步优化时空特征
- 结合神经微分方程建模平滑相机轨迹
- 适合需要实时渲染的动态场景应用
高效神经视频表示对视频压缩到交互式模拟等应用至关重要。现有方法常面临内存占用高、训练时间长和时序不一致的问题。本文提出一种新型神经视频表示,结合3D高斯点云与连续相机运动建模。通过神经微分方程,该方法学习平滑相机轨迹,同时以高斯点保持显式三维场景表示。此外,引入时空分层学习策略,逐步精炼空间与时间特征,提升重建质量并加速收敛。该内存高效的方案实现了高质量快速渲染。实验表明,分层学习与鲁棒相机建模相结合,在多种视频数据集上均取得先进性能,涵盖高运动与低运动场景,具备强时序一致性。
原文摘要 · Abstract (English)
Efficient neural representations for dynamic video scenes are critical for applications ranging from video compression to interactive simulations. Yet, existing methods often face challenges related to high memory usage, lengthy training times, and temporal consistency. To address these issues, we introduce a novel neural video representation that combines 3D Gaussian splatting with continuous camera motion modeling. By leveraging Neural ODEs, our approach learns smooth camera trajectories while maintaining an explicit 3D scene representation through Gaussians. Additionally, we introduce a spatiotemporal hierarchical learning strategy, progressively refining spatial and temporal features to enhance reconstruction quality and accelerate convergence. This memory-efficient approach achieves high-quality rendering at impressive speeds. Experimental results show that our hierarchical learning, combined with robust camera motion modeling, captures complex dynamic scenes with strong temporal consistency, achieving state-of-the-art performance across diverse video datasets in both high- and low-motion scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。