用深度模型动态过滤干扰,提升单目视觉定位精度
Dynamic Visual SLAM using a General 3D Prior
- 结合深度预测与光斑优化,实时分离动态与静态区域
- 在动态场景中实现更稳定的相机位姿估计,显著减少漂移
- 适合机器人导航、AR/VR等需高鲁棒性的实时定位场景
可靠地增量式估计相机位姿与三维重建是机器人、交互可视化和增强现实等应用的关键。然而,在存在动态物体的自然环境中,场景变化会严重降低位姿估计的准确性。本文提出一种新型单目视觉SLAM系统,可在动态场景中稳健估计相机位姿。通过结合基于几何光斑的在线捆绑调整与最新的前馈重建模型,我们设计了一个前馈重建模型,可精确滤除动态区域,并利用其深度预测提升光斑式SLAM的鲁棒性。通过将深度预测与捆绑调整估计的光斑对齐,有效缓解了前馈模型批处理带来的固有尺度模糊问题。
原文摘要 · Abstract (English)
Reliable incremental estimation of camera poses and 3D reconstruction is key to enable various applications including robotics, interactive visualization, and augmented reality. However, this task is particularly challenging in dynamic natural environments, where scene dynamics can severely deteriorate camera pose estimation accuracy. In this work, we propose a novel monocular visual SLAM system that can robustly estimate camera poses in dynamic scenes. To this end, we leverage the complementary strengths of geometric patch-based online bundle adjustment and recent feed-forward reconstruction models. Specifically, we propose a feed-forward reconstruction model to precisely filter out dynamic regions, while also utilizing its depth prediction to enhance the robustness of the patch-based visual SLAM. By aligning depth prediction with estimated patches from bundle adjustment, we robustly handle the inherent scale ambiguities of the batch-wise application of the feed-forward reconstruction model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。