无需已知相机位姿,可处理任意长视频的3D高斯点渲染方法
Towards Better Robustness: Pose-Free 3D Gaussian Splatting for Arbitrarily Long Videos
- 利用视频帧间连续性设计相邻位姿跟踪,稳定估计相机姿态
- 通过高斯可见性保留检查,自动分割长视频并分段优化
- 在多个数据集上优于现有方法,适合长时序视频重建场景
3D高斯点渲染(3DGS)因其高效与高质量渲染成为主流表示方法,但其训练需已知每张输入图像的相机位姿,通常由运动恢复结构(SfM)流程获取。现有工作尝试放宽此限制,但在复杂轨迹的长序列视频中仍面临困难。本文提出鲁棒性更强的Rob-GS框架,可逐步估计相机位姿并优化3DGS以处理任意长度视频。通过利用视频固有的时间连续性,设计相邻位姿跟踪机制,确保连续帧间姿态估计稳定;为应对任意长输入,提出高斯可见性保留检查策略,自适应将视频序列划分为多个片段并分别优化。在Tanks and Temples、ScanNet及自采数据集上的大量实验表明,Rob-GS性能超越当前最优方法。
原文摘要 · Abstract (English)
3D Gaussian Splatting (3DGS) has emerged as a powerful representation due to its efficiency and high-fidelity rendering. 3DGS training requires a known camera pose for each input view, typically obtained by Structure-from-Motion (SfM) pipelines. Pioneering works have attempted to relax this restriction but still face difficulties when handling long sequences with complex camera trajectories. In this paper, we propose Rob-GS, a robust framework to progressively estimate camera poses and optimize 3DGS for arbitrarily long video inputs. In particular, by leveraging the inherent continuity of videos, we design an adjacent pose tracking method to ensure stable pose estimation between consecutive frames. To handle arbitrarily long inputs, we propose a Gaussian visibility retention check strategy to adaptively split the video sequence into several segments and optimize them separately. Extensive experiments on Tanks and Temples, ScanNet, and a self-captured dataset show that Rob-GS outperforms the state-of-the-arts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。