arXiv:2501.01949cs.CV2025-01被引 10

用分块分层对齐法,快速把单目视频变3D场景。

VideoLifter: Lifting Videos to 3D with Fast Hierarchical Stereo Alignment

  • 分块局部用可学习3D先验注册片段,提升效率
  • 全局用树状结构合并,减少82%训练时间
  • 适合需要高速3D重建的VR与机器人应用

从单目视频高效重建3D场景仍是计算机视觉的核心挑战,对虚拟现实、机器人和场景理解至关重要。现有逐帧渐进式重建方法无需相机位姿,但计算开销大,长视频时误差累积严重。为此,我们提出VideoLifter,一种基于片段的局部到全局策略的视频转3D新框架,实现极高效率与当前最优质量。局部层面,利用可学习3D先验对片段进行注册,提取关键信息以初始化3D高斯,并强制片段间一致性,优化效率;全局层面,采用基于树的分层合并方法,通过关键帧引导进行片段对齐,成对合并并结合高斯点剪枝,再联合优化,确保全局一致性的同时有效缓解累积误差。该方法显著加速重建过程,训练时间减少超过82%,且视觉质量优于现有最先进方法。

原文摘要 · Abstract (English)

Efficiently reconstructing 3D scenes from monocular video remains a core challenge in computer vision, vital for applications in virtual reality, robotics, and scene understanding. Recently, frame-by-frame progressive reconstruction without camera poses is commonly adopted, incurring high computational overhead and compounding errors when scaling to longer videos. To overcome these issues, we introduce VideoLifter, a novel video-to-3D pipeline that leverages a local-to-global strategy on a fragment basis, achieving both extreme efficiency and SOTA quality. Locally, VideoLifter leverages learnable 3D priors to register fragments, extracting essential information for subsequent 3D Gaussian initialization with enforced inter-fragment consistency and optimized efficiency. Globally, it employs a tree-based hierarchical merging method with key frame guidance for inter-fragment alignment, pairwise merging with Gaussian point pruning, and subsequent joint optimization to ensure global consistency while efficiently mitigating cumulative errors. This approach significantly accelerates the reconstruction process, reducing training time by over 82% while holding better visual quality than current SOTA methods.

3D重建单目视频高效算法高斯渲染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。