arXiv:2412.04463cs.CV2024-12CVPR被引 243

用深度视觉SLAM框架实现动态视频中快速精准的相机与深度估计

MegaSaM: Accurate, Fast, and Robust Structure and Motion from Casual Dynamic Videos

  • 改进训练与推理策略,使深度SLAM适用于复杂动态场景
  • 在真实与合成视频上精度显著优于已有方法,速度更快或相当
  • 适合无固定视角、小视差的日常拍摄视频处理

我们提出一个系统,可从随意拍摄的单目动态视频中准确、快速、鲁棒地估计相机参数和深度图。传统结构光与单目SLAM方法通常假设输入视频以静态场景为主且具有大视差,此类方法在缺乏上述条件时易产生错误结果。近期基于神经网络的方法虽尝试克服此问题,但往往计算开销大或对未知视场角及非受控相机运动表现脆弱。我们证明了深度视觉SLAM框架经适当改进后,在复杂动态场景、无约束相机路径下仍能有效运行,包括视差极小的视频。在合成与真实视频上的大量实验表明,本系统在相机位姿与深度估计上显著优于先前及同期工作,且运行时间更快或相当。交互式结果见项目页:https://mega-sam.github.io/

原文摘要 · Abstract (English)

We present a system that allows for accurate, fast, and robust estimation of camera parameters and depth maps from casual monocular videos of dynamic scenes. Most conventional structure from motion and monocular SLAM techniques assume input videos that feature predominantly static scenes with large amounts of parallax. Such methods tend to produce erroneous estimates in the absence of these conditions. Recent neural network-based approaches attempt to overcome these challenges; however, such methods are either computationally expensive or brittle when run on dynamic videos with uncontrolled camera motion or unknown field of view. We demonstrate the surprising effectiveness of a deep visual SLAM framework: with careful modifications to its training and inference schemes, this system can scale to real-world videos of complex dynamic scenes with unconstrained camera paths, including videos with little camera parallax. Extensive experiments on both synthetic and real videos demonstrate that our system is significantly more accurate and robust at camera pose and depth estimation when compared with prior and concurrent work, with faster or comparable running times. See interactive results on our project page: https://mega-sam.github.io/

视觉里程计深度估计动态场景单目SLAM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。